Pith. sign in

Paper Citation Record · LEDGER

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

As of 17 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2506.23283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23283 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:59.930418Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact4
  • verified fuzzy43
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0566584-8b1a-4e89-bc30-b1d8a7a49a6e · outbound

This paper cites Vivit: A video vision transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.381794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.381794Z digest=sha256:4f38cfef258067fc0129bd46e8da9db3dc56cb68239b3196a8d0da083a255034

Observation 18a8fd3e-b682-462a-913c-6f3ea3a4878d · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition BEiT: BERT Pre-Training of Image Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.473450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.473450Z digest=sha256:8629830efd9923a499af0bcce65100bfe126ffe71a04f4d24baa12708941ad93

Observation 97822a62-148e-4a60-ad27-4c0290eedf20 · outbound

This paper cites Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.566560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.566560Z digest=sha256:a1929c6f8df73e18a355fbdcd568a7234fec6765f9e73613b6037c7b37b5ad59

Observation 0dd5bef8-715c-41db-a180-87c6bea1eb99 · outbound

This paper cites Coyo-700m: Image-text pair dataset, 2022.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coyo-700m: Image-text pair dataset, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.684318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.684318Z digest=sha256:7bf460b04312b95557b2bfeb78698c68a0b7f0f1eb71d8b42639d5577b23724b

Observation 066d303e-31f0-495a-a9d5-e53f30ab3fcf · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Quo vadis, ac- tion recognition? a new model and the kinetics dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.823473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.823473Z digest=sha256:14a68f564a7c2b52188ca1692b7ca7ac2b29d45f7219067ca0933f5d0f6a5373

Observation 6b1b048b-180e-4829-9f5b-7141258a6a96 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.963310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.963310Z digest=sha256:c9107dd6256a21fe55d2508c5a496e26e518feea2f04ee5099acbdcc9d94fbb0

Observation 745bc955-cd20-424a-bb12-2e2c8673636e · outbound

This paper cites A simple framework for con- trastive learning of visual representations, 2020.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition A simple framework for con- trastive learning of visual representations, 2020

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.088191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.088191Z digest=sha256:281f2ef79568e3fb090ab75d5d336391f68045c9f8a9921b0ec93c3b715c6a2d

Observation 35f85598-e66d-4960-b93e-a9b3d5b956d3 · outbound

This paper cites Feature-wise transformations.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Feature-wise transformations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.215965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.215965Z digest=sha256:83e6678b73466d574541afacda37f166ed8bc6e8ca37162ed2667505482ed687

Observation ae2ebe95-7d61-4550-8588-d8c48b4c95d4 · outbound

This paper cites Multiscale vision transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiscale vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.324371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.324371Z digest=sha256:73ccb564a0b016cf595c87e2b9178cfce99d5c7fb2c93c2b527b1339839393cc

Observation 56401364-3da8-488c-ad07-e74ededb08f6 · outbound

This paper cites X3d: Expanding architec- tures for efficient video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition X3d: Expanding architec- tures for efficient video recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.451686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.451686Z digest=sha256:88faa56d0b840731cc260f334331cb0df3fc35674aedd8a4998df79d93b50ef1

Observation 2cfc5cce-ae53-43a3-9dfa-93799361f043 · outbound

This paper cites Masked Autoencoders As Spatiotemporal Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Masked Autoencoders As Spatiotemporal Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.583049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.583049Z digest=sha256:7001de2c5f31fd83f31bc699dd6a51c51e7e16dec6fb6438d2c3072f33fd85e0

Observation c259ede9-b86a-4879-9f91-ae3ec10321ff · outbound

This paper cites Slowfast networks for video recog- nition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Slowfast networks for video recog- nition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.697762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.697762Z digest=sha256:5e352096a585522258355e5c867df689680ed9f41ea34b05d2c52577a9221528

Observation 5da590e6-82da-4603-8c2e-6b59e9ac8884 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The” something something” video database for learning and evaluating visual common sense

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.327642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:51.830505Z digest=sha256:a67778fae1afa1a6b9c2432df966e8e00fb7f344ca5a0ecf827c864d569403fb

Observation 68690598-06ab-4175-870e-f52779ee66e8 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.926035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.926035Z digest=sha256:e1c82aee8c4a7184acc63f88e24ba6b222da330bebfa7fa25b9279a1bb51707f

Observation 39b35f9b-124f-4e9d-90f8-3ea5f91027d2 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Efficiently Modeling Long Sequences with Structured State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.045734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.045734Z digest=sha256:1b91f2ac3964f2c65f8a8eed939b1616531432ae457ff4ddc5d83e2a96a0d6f7

Observation 70f67fe2-2691-497b-a62c-b808bbe5a4b0 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary object detection via vision and language knowledge distillation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.279929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.127801Z digest=sha256:58d443c5a1faba54059c99b7064de0a7ea65075b5072ca22aab962d4ffaa72bc

Observation d3fe1164-933b-4afb-abab-edf756748487 · outbound

This paper cites Trustworthy machine learning: From data to models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Trustworthy machine learning: From data to models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.195299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.256779Z digest=sha256:e868f5f64b339f4ddf15d78227d14ff166966987e6d20382d1476e02fe0d7908

Observation 245cfa19-8e54-4bed-a6b5-431633cdeba3 · outbound

This paper cites Turbo training with token dropout.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Turbo training with token dropout

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.084240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.390034Z digest=sha256:72da482df3167334ddfd538600ec7b8f25e48b31e07f6ce481e3fbbf38f98a9b

Observation 75dbef3c-98ee-4f7e-9e9e-9b8c55438318 · outbound

This paper cites Learning spatio-temporal features with 3d residual networks for action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning spatio-temporal features with 3d residual networks for action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.978319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.549772Z digest=sha256:606175e9d59839de6dd88d6bb9c497f754a81507b8f664203b04d121c2a673b4

Observation fd048668-ac8d-4cbe-87be-bcb5b743aee7 · outbound

This paper cites MambaVision: A Hybrid Mamba-Transformer Vision Backbone.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.633842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.633842Z digest=sha256:d3063dbba143e9f2942cfc33ccd78ee64ae3b6c4429be38f6fdc6b1c1a9bdea2

Observation a4995629-e6d5-4ab9-a680-a5d3da6ba861 · outbound

This paper cites Clipscore: A reference- free evaluation metric for image captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Clipscore: A reference- free evaluation metric for image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.875378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.745648Z digest=sha256:c1415996be8e911de894992941c167ad4b411b24a8a3635556f445ab72000c8e

Observation afba095f-b5d7-4a4d-a56f-5b12f2cf32f8 · outbound

This paper cites Arbitrary style trans- fer in real-time with adaptive instance normalization.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Arbitrary style trans- fer in real-time with adaptive instance normalization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.750886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:52.831996Z digest=sha256:af0a6d6b3cf0660de793f0cdda0e98387ac66c830ae2e38e347c01e5b4960255

Observation 250059a2-427d-4720-9901-4c9cb5b8384e · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.943841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.943841Z digest=sha256:77d73d864fbaf45abc61e0296cd3a55ca7bdd1fe70add00719c5512254b630a5

Observation 75807e7a-5368-46d2-982f-815c6503e810 · outbound

This paper cites Smeulders.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Smeulders

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.626186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.022920Z digest=sha256:8531ce80565d5bb52d18442a0c5f8834700287573fdda5bcdbbd73605aa00c79

Observation 5c017a9e-706f-40cf-af86-3ef5a43b02f3 · outbound

This paper cites Long movie clip classification with state-space video mod- els.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Long movie clip classification with state-space video mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.504852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.133192Z digest=sha256:031e20b340f4502dbf15c0ff30c680b708d828ce8d8b43e4d3b090537bb0cd4f

Observation 034efcb6-bfc0-4177-9cbf-e9149da6fe12 · outbound

This paper cites Scaling up visual and vision- language representation learning with noisy text su- pervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling up visual and vision- language representation learning with noisy text su- pervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.396778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.286576Z digest=sha256:81874657b4522fcf97f73309a36397e73753c300551f74abad7bb8a2db5a44cc

Observation 06acb9ba-58e8-401a-b5ea-eddbfbff387e · outbound

This paper cites Laine, and Timo Aila.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Laine, and Timo Aila

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.285784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.420848Z digest=sha256:3aa1ad8a72c44830ac7df01eceb33d951cdd1e1fa569f039f1a74783707838f7

Observation c231f89a-87ef-43fe-b2bd-7cd483c811f4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The Kinetics Human Action Video Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.493512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.493512Z digest=sha256:f1ac6d59d9081dec91769f7a5f16f9feb62119f8aa80c8a1ad44b29f1892b904

Observation 4e4c4956-8e4e-498e-9b53-449f2cedceab · outbound

This paper cites Kuehne, H.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Kuehne, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.170200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.561748Z digest=sha256:20ef61047629ce766cba7bfd3d95201692fa4358bd1fc39feaf3e7c634baa193

Observation 8821f97d-59ee-4773-9bee-8fc6a466d9c6 · outbound

This paper cites The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.036537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.693704Z digest=sha256:58da56e48e55e4016f306055e1fd7a1160aa050a692b044f067ce69151bc962f

Observation 11c6d7db-7c34-4bcf-985a-a185bbf26bb1 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.789728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.789728Z digest=sha256:53744b827bedc794522e919ffa38a06487327e9e150713b2ed029dc2f95700e1

Observation 965818a5-d127-4beb-b212-05a651f0ad89 · outbound

This paper cites Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:01.162745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:53.897777Z digest=sha256:ef2101a625cf95badcf82c3e3ae7bbf5fc40742063509b2c45f7507115a155c2

Observation 6c983faf-b075-426a-beb7-3bd8197b733e · outbound

This paper cites Videomamba: State space model for efficient video understanding, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Videomamba: State space model for efficient video understanding, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.830280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.013868Z digest=sha256:77fedff9075d59e921835f5ee84fb1be44d73b2209a2793f95e23dab4d8c6d29

Observation 02fcb1cf-32a1-4da0-b55e-7c84cb0018dc · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unmasked teacher: Towards training-efficient video foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.653523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.127628Z digest=sha256:0e30f8978a54bb5247e4d114e2c13bc6482f68ea776d0ebf54aff424d674c134

Observation 27f5fb1b-f812-4cce-97ce-0a366490a1c7 · outbound

This paper cites Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.465539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.238037Z digest=sha256:bfd4e8d45938f1f46eb72ec486d1e56a35578955a7111d62e5e0e270edff352e

Observation 26abb7e7-33e8-4293-b6cf-5b079c84c844 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.330276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.367511Z digest=sha256:131121e4854ac374ad721b1f0db2d4f791521cb22a8f52b01407a530ae91adac

Observation 1141e522-dbbe-4950-b689-f4c2bf95f6ed · outbound

This paper cites Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.141385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.431653Z digest=sha256:a3cef9e0dee08f09c06953fb3cdd8ce7b092a40650124b1738bc89b938bc9be4

Observation 83f7f27f-ffb3-4433-80db-6750c1eb9ce2 · outbound

This paper cites Pointmamba: A simple state space model for point cloud analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Pointmamba: A simple state space model for point cloud analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.960882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.524479Z digest=sha256:704abf9b98a93d4f5c78777036bcef0b02c88ee6aa82081437d39885c61a5752

Observation 096c446c-94c6-459f-8729-50371225bc84 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Jamba: A Hybrid Transformer-Mamba Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.612897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.612897Z digest=sha256:e9de6aaaf3571e297fa9e8e5e5f1ff0bb503f70ba8a5525b5e71497b0531babf

Observation e29d25ed-702c-46a4-ac21-893c111fc511 · outbound

This paper cites Learning to recognize procedural activities with dis- tant supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning to recognize procedural activities with dis- tant supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.822263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:54.696571Z digest=sha256:e15f409f909a8ec4b546a4859779c9cd9692e6bbd64d6afe09a9a7d23b249c1c

Observation 82ff3b0f-f212-4e6f-88d6-b2b3490609bf · outbound

This paper cites Frozen CLIP Models are Efficient Video Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Frozen CLIP Models are Efficient Video Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.788646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.788646Z digest=sha256:3d5d910f563bdc4692c3793b4b64a75ec7076eacb0eb7f0bb3c3955ac26ef022

Observation 9898fcdc-ced0-4e8e-9241-e170e65e6353 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Annotation-free Audio-Visual Segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.889200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.889200Z digest=sha256:78696bb0bc35fc746d70d74a05ff6d562a8069c01337cf7cd7014542bf766657

Observation b070c1c1-bf0e-4edd-b688-d0623e3d3f36 · outbound

This paper cites VMamba: Visual State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VMamba: Visual State Space Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.973660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.973660Z digest=sha256:8b319088ded4535c571df5fa0b6138a6a5a85552e468d003e8096d824084a168

Observation 3b47763b-25ec-4a8e-bbc0-c867acec771c · outbound

This paper cites Video swin transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video swin transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.670936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.081355Z digest=sha256:5ee673f3206a4b865fe0e15a290b8f97c56d5bd702c74d9bbbb55ff9ffc6f444

Observation 6170fa3d-63f0-4201-b1f0-b538fecb1aa3 · outbound

This paper cites Freesegdiff: Annotation-free saliency segmentation with diffusion models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Freesegdiff: Annotation-free saliency segmentation with diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.524175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.165046Z digest=sha256:f1ee97f35b5368dce465f9e6121c7d2f280cb16687095c53f89960fc85df5a19

Observation 03621015-443a-4b08-92ce-7b5cb1f70656 · outbound

This paper cites DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.235358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.235358Z digest=sha256:eb5936c9447af0d5b4debb1dddb650500a065843116ad11a25942bac7fc8005e

Observation e52b8b84-b3f9-4342-bf72-92861215bdd6 · outbound

This paper cites Open-vocabulary semantic segmenta- tion with frozen vision-language models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary semantic segmenta- tion with frozen vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.336866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.327089Z digest=sha256:214cb09ba828114cffe7fbd1ab489b4dc82c3b8b4325985eaa6e6a1c1c90e1f3

Observation 69c49ab9-c49e-4727-87ca-5df1356d7e02 · outbound

This paper cites Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.120762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.420685Z digest=sha256:3c1cc26a794148e6556a767cb89f6b2a6be285ef5ccbca6583c5a0e75ddb3a94

Observation 2573ff8e-3f54-4079-b207-f54ad18e4c47 · outbound

This paper cites Scaling open-vocabulary object detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:07.890235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.482406Z digest=sha256:c5453d32d8d71c9766f1d7230b6f0a82a0803878f3754f40e83469248b6f45be

Observation 46177889-2581-457d-bd97-0e86f0a016ed · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:c9d7d055fe71afafe16919c7725bd1a434954bdc8d3216b608cea16c13b8ca16

Observation 8bad67b6-8264-424e-98ae-4d4ea9723705 · outbound

This paper cites Expanding Language-Image Pretrained Models for General Video Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Expanding Language-Image Pretrained Models for General Video Recognition

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.859981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.663136Z digest=sha256:9c2b911da2a2180b9c911051e4876a68535575ff4ddc4d6a3f5b0be0c1c723ad

Observation 00a9a893-7320-4bb5-884a-67d9c6c1167c · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.746646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.746646Z digest=sha256:cc18a6e9dba7f4752e864cc42f8f4e1d355e1f65577b9af60c7caba53d576cd8

Observation 1fceb07c-de6f-4b34-9b3b-72af3549bc11 · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:07.162702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.855563Z digest=sha256:5490a2a37d247b43b9323d2780506983b4ecff87d57bef83d8be488ff45e7e52

Observation 600a5af7-0e65-45fa-9ef4-dc11436ed1f6 · outbound

This paper cites ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.638728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:55.932723Z digest=sha256:aa17bb2ec7fd7a7505eb10dbd22468044e0ab253c5f153634135c01f6f7b6d9a

Observation 47ad84a8-cfe1-489d-8460-5ab7d35ae11f · outbound

This paper cites MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.425970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.065967Z digest=sha256:51b5cdee14a4ce68fdbd0e48388c2de7d7309e71a579778779599009e6815c7d

Observation 01124fa0-a327-49ef-ac19-31cb7f5beac9 · outbound

This paper cites Dual-path adaptation from image to video transform- ers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dual-path adaptation from image to video transform- ers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.981489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.170231Z digest=sha256:76857c91c2d79fdedfe2ba121ad7e48c0fee376b16b72d40206c162820abf2d3

Observation 56679b2e-1c87-4056-a665-fe6bc03ccdb6 · outbound

This paper cites Peebles and Saining Xie.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Peebles and Saining Xie

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.766762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.291434Z digest=sha256:c54256b148cdd0a20362fc27847155c6ea9672ef19d95f3eb5503ca45c3c6f95

Observation e3362da8-a59f-4667-bb1e-fe3b44df7dbd · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:06.501535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.418109Z digest=sha256:718e321a3b3b0301c44b4f681c93f0a2793d9512421320114418b7fb7c5c58ad

Observation 09d3aab8-921b-417a-ba0d-7e92b0719a3c · outbound

This paper cites Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.295700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.546195Z digest=sha256:0b1077d8a1fc60e258b328907b77d5b39dd6158d61853fd4a1e72f2c9d3292a8

Observation 4c805022-7f59-4adb-be59-353adf757540 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning transferable visual models from natural lan- guage supervision

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.119792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.659418Z digest=sha256:6b650fd84e46ffc16d8c3cd18f808dd8d51a094b79ea58b2d5a14dd58f76c365

Observation ead28b01-3e89-46be-8e2f-ef446ded4382 · outbound

This paper cites Token- learner: Adaptive space-time tokenization for videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Token- learner: Adaptive space-time tokenization for videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.975865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.759866Z digest=sha256:6087354e53644641b4db191c99f7d3f2c5c16a890fc609ed9f76531c9f1a2c4e

Observation 7c17898b-796b-46d0-afdd-8836a651f2b1 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:56.863163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:56.863163Z digest=sha256:a24b0318ce734d7c067435ee138b2366a87c545d7749b8d8e97c6c701b2af854

Observation 8b354a31-49bd-4357-9a09-10a6cd7e5e34 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Only time can tell: Discovering temporal data for temporal modeling

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.790063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:56.962660Z digest=sha256:3afdd574dc559438cc36de27bc887fd59e9796b0c5314bdde13ae8184bd67cdb

Observation 8c0e1250-f422-4fe9-9601-1a08ccdc984d · outbound

This paper cites Darf: Depth- aware generalizable neural radiance field.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Darf: Depth- aware generalizable neural radiance field

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.620347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:57.082523Z digest=sha256:d584be64bc1b3c05a14b6a273b7baa67c4ede9461a2dca21b82a409cdd305192

Observation 28917599-655e-4971-8114-058987c3fdcb · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.234563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.234563Z digest=sha256:301b131614ea428f816d0b0764188ddd32ca15ea78e7a9a2945d0b63b4ec3654

Observation 779064c5-4c2f-41b4-a898-307d77d7c2ca · outbound

This paper cites Coin: A large-scale dataset for comprehensive instruc- tional video analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coin: A large-scale dataset for comprehensive instruc- tional video analysis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.432865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:57.360492Z digest=sha256:1db96db24cc5e617ba55aace0100f149d0144c8faaa9f5be3789c668cafac20b

Observation fbdc6a50-7eec-48f7-8c45-b34de633b22f · outbound

This paper cites Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.938654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:57.464332Z digest=sha256:6b0d693ead976f4c85d1821f166dd3e6fda36e08df3e5d12205859502160131e

Observation af449d3e-434b-4f99-ada6-63a5585c1387 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.583666Z digest=sha256:444c08f298c8845c721a37d39736fa119b861bc44fc712c4552024bde6c61e46

Observation 0c40be8d-6286-4410-9506-79a9f25f2837 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition An Empirical Study of Mamba-based Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.712328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.712328Z digest=sha256:a9c5e71e302423546bcb7577042d9e80d0461b1d2019dd0473478855b44ba55b

Observation a3092b77-db40-4209-bc01-513dabb0a114 · outbound

This paper cites Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.852368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.852368Z digest=sha256:c59d010b600b5f2d094ea1dad7f47fcfb814974c2d58f666c0079f589dd5a8be

Observation fefb174e-7ad2-4264-838b-62483d589634 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.921969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.921969Z digest=sha256:fb3711cd94e66262a7bc2539c680c9e784472db8d278a673f78475823e7d1841

Observation 897aba94-bd40-4570-a160-c81813c23320 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.007202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.007202Z digest=sha256:c3c27f7832ecb4a8fb23717b0c28d91cc15af0a51a05cf85bb3ed60408c82276

Observation 6cdd6186-d78a-4fba-8888-cf116c13b917 · outbound

This paper cites PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.110217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.110217Z digest=sha256:d897a2a3d4b1f3c284bfb0d92f931c3eb826c6009623e8687771ee145c5b9730

Observation 94f0fbf7-8e97-43b8-a1e3-e50cb9b47a84 · outbound

This paper cites DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.193805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.193805Z digest=sha256:0fc74496ba91cda98ab9422f1f9163976703c993018c893fc6ab99c47d498cbe

Observation 7d526405-4e77-4079-a62e-68b17786e9ee · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.251948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.251948Z digest=sha256:9a207ea3ebf04bc1f171527be8f89e1e328ba07ffd75e1a2d15b2e8aee3d9d2d

Observation 90e734f1-29e5-4dc2-8f66-3a7369f14b76 · outbound

This paper cites Multiview transformers for video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiview transformers for video recognition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.304085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:58.324442Z digest=sha256:1910a883da50b6a81a0c18186c04c4047a11e002ee08d826c948756b5011838f

Observation 68feddfa-c382-4fae-bb7d-49084f20ac09 · outbound

This paper cites MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.389943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.389943Z digest=sha256:680f339c47fcadfccc2b9665582a8acb32c0734b7893743586320a8e3b040439

Observation f2ed0136-122a-4a41-902c-67b430d24944 · outbound

This paper cites AIM: Adapting image mod- els for efficient video action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition AIM: Adapting image mod- els for efficient video action recognition

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.146644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:58.460734Z digest=sha256:07c83b696994c4f0269285e2f1a4b7d6552194da4df7241e83e7f659156a1ec1

Observation 8c6e1585-036c-4a36-987c-68a3ef9dc655 · outbound

This paper cites Multi- modal prototypes for open-world semantic segmen- tation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multi- modal prototypes for open-world semantic segmen- tation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.959704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:58.591097Z digest=sha256:8e21fc0e679acba3946ffd44c501d8a35090b33af955dd909e6cb6595a55613b

Observation f07b56a4-579a-47ed-ad1b-744a7fafac25 · outbound

This paper cites Remamber: Referring image segmentation with mamba twister.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Remamber: Referring image segmentation with mamba twister

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.772877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:58.763767Z digest=sha256:39e58abbe1b3c44c3c81e6eabe0bc367038a1b0709d730cfaaf77bdcde0abcc2

Observation c229ec09-62d6-4d04-b8e3-a67848f1fee9 · outbound

This paper cites Learning with multi- class auc: Theory and algorithms.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning with multi- class auc: Theory and algorithms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.585717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:58.912661Z digest=sha256:2106a5283e46e5363efc8260ecce2eeaaf943f7007bb5b01a3fd1bfc15f1a365

Observation 9a409808-e05f-4f5c-a661-b69f11de62a6 · outbound

This paper cites Optimizing two-way partial auc with an end-to-end framework.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Optimizing two-way partial auc with an end-to-end framework

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.417014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:59.152808Z digest=sha256:e0add720c5c4be776c3742a8e3035bbaf02c05c11150d50915b442724079cfd7

Observation 7cbfadaf-943e-4673-8f37-e4d2aad83fb1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigmoid Loss for Language Image Pre-Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.342925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.342925Z digest=sha256:db5b35cc68c808a4979d6582905f8a74c8f580ce8aa4f14aa1cf1f4b2f7a4147

Observation c448c7b1-0717-4d7e-baff-928868910d1e · outbound

This paper cites Point Cloud Mamba: Point Cloud Learning via State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Point Cloud Mamba: Point Cloud Learning via State Space Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.480981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.480981Z digest=sha256:8efdb956444d78efca02a6b6eb63e6500630fa40fb952f574cb5030e1297bb07

Observation 613a4a46-2840-47e1-a6a4-eb0402e8d530 · outbound

This paper cites G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.564079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.564079Z digest=sha256:f526df07f10ae0af79ec5729ea67d1797712c8c48918b5e44f6d64659e280c7c

Observation c101d2c5-237a-4331-8e65-f43b9565ab7e · outbound

This paper cites Vidtr: Video transformer without convo- lutions.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vidtr: Video transformer without convo- lutions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.118114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:59.662077Z digest=sha256:97e10afeedb437bb8512b43e9cf68171e1a102046ea06ca2e33e20cbd798a985

Observation 8014a4aa-3a2c-4def-8347-74f7b1beac90 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Regionclip: Region-based language-image pretraining

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.837632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:59.733742Z digest=sha256:f7d011a5aa8ed4e1a7bc79813e0562cf99a5b5ed285fe281e0b97f306d604b05

Observation 6deeac2c-3a55-4d5d-8cb7-97eb4bc75e57 · outbound

This paper cites Graph-based high-order relation modeling for long-term action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Graph-based high-order relation modeling for long-term action recognition

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.571641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T21:51:59.857627Z digest=sha256:4a74cd4b5c71f8eeac3e7f630e39e1e6983a9288701589b6d1c73fc918d99eea

Observation ce255339-f313-4b06-8e78-b21533f4e672 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.930418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.930418Z digest=sha256:437cccb2caf4b3e06164981e0ace48ea1ba52172c0285ea60b5daf89d0ac9e2d

Pith citing papers

No inbound Pith citation observations are available.