Pith. sign in

Paper Citation Record · LEDGER

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 6 inbound Pith citation observations for arXiv:2506.10915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10915 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:21:19.076614Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:47.529772Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:44:59.794530Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7e0ff73-8461-4a52-aea9-cda721dab9a0 · outbound

This paper cites Pika art.https://pika.art, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pika art.https://pika.art, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.782774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.766104Z digest=sha256:75b63cfe1cbbc733bbdfd6c7719c0264c4192fb67d35ada7c4431c547d5679cb

Observation 02eda61e-eb01-498f-8d63-e897e9cd6f46 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.818385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.818385Z digest=sha256:7c4d63543f455aff11b43ed7c49fe4792f579693b4c4dac42b2b14218b66b8f4

Observation 99b9b93a-2590-49e7-8816-6e57030e277d · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2023.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flux.https://github.com/black-forest-labs/flux, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.769840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.821935Z digest=sha256:42837279b73f70865e86dc5dc3aa39d3db5163099af20e0a56d4a3a5de635698

Observation 60d7f820-24d1-4853-af12-8bb281537e84 · outbound

This paper cites Video generation models as world simulators.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.923561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.923561Z digest=sha256:43153eca2a5eb38c5ccc41a3a7aa3b1e32f18a643b5859a791cb7b893c497abc

Observation be5e0b4a-7155-47a8-a2dc-e306aed9dd50 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.927125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.927125Z digest=sha256:58a682d503116e72a8ff45b95801262434ec2601af5bcf519dc09093a56d3443

Observation ca55e251-000c-4ef7-aa51-3fa580c03782 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.934563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.934563Z digest=sha256:956b1c8c54d406c578c0aefb7fb9e7607059e7f8a6961dc4996ca36c9eed84f1

Observation b2836212-e280-4e41-aa74-47852d59b0ec · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Goku: Flow Based Video Generative Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.937837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.937837Z digest=sha256:c9d33620e5c639582f204cc28ef34028695eac692e1fcf576d3f011124e0004d

Observation d86eef98-54df-4310-9284-7862710e3e4c · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.940770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.940770Z digest=sha256:7a17c0479075c7e8c6b84a63b216dc642139671ae87baf471e9edfec5ea0290a

Observation d2dbefb8-b6ce-4d60-add6-3c613861d226 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.943726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.943726Z digest=sha256:af8de3a3f86112b596010b09ae1ee2bc97fb694b2dd1116e5a526cff7ceab419

Observation 52f37071-c193-4085-8050-6f41ff3fbf83 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.946935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.946935Z digest=sha256:907b7cf8e4babcd8a9501fa5926fa2f9f9cba6d3f4eb6cfb01b2cc6e8fdf0020

Observation e913141b-a757-4a09-a3fb-9a00293ffc7d · outbound

This paper cites Dimba: Transformer-Mamba Diffusion Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Dimba: Transformer-Mamba Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.950027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.950027Z digest=sha256:a00b3c64afc507ba5c4a6a3fe8e3c2964c5d9b2e126a8aa97ffd36cbd460606d

Observation 47d6df64-ca99-4439-8328-8a0ffc12025e · outbound

This paper cites LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.953995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.953995Z digest=sha256:9c631655af97f7130de77870609fdba22e61b608a5b6588b4a6329714200badd

Observation ab791d2d-c81f-4b35-962c-1a12e938b4fa · outbound

This paper cites Matten: Video Generation with Mamba-Attention.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Matten: Video Generation with Mamba-Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.957450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.957450Z digest=sha256:5328535fe41d90b20885dc55707972b07fbc0d5ed69708c2c9217c686780fbcd

Observation d0462d30-c1b8-459d-be5f-85656edbcbc0 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.960505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.960505Z digest=sha256:669c0455d01fc7599ec977db0de83477b58ee6a660ced402400e73a7506f5a25

Observation 27f80f9b-8858-4905-8597-1a94c3ee5496 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Efficiently Modeling Long Sequences with Structured State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.963248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.963248Z digest=sha256:e9eb6661a00cc57c3ae84a2b265af6a0d246ca7572fe31bc6c9a346cfbea4e1d

Observation 41e58600-93a5-434e-b8c8-c10c9be8c886 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.965955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.965955Z digest=sha256:ea9d5ec9cf751683ffa2f5b5ad303b407f7359de3fb46970c4a8d4a57205b7c8

Observation 21ac86b8-ebd5-4013-ba11-9b4cc914480f · outbound

This paper cites Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.968467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.968467Z digest=sha256:8fff4fb931012267aa80f9560d8ef68ad36b481e38a8bebd1a7bc6d28bb5f87f

Observation 1989c988-afa6-4980-a0f8-6a9d576cd81c · outbound

This paper cites Zigma: A dit-style zigzag mamba diffusion model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Zigma: A dit-style zigzag mamba diffusion model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.735860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.970874Z digest=sha256:0cf8caedd8b8b878cfbb8ed820573858f5ce5113df9a63979f4c334723e97e14

Observation 541e0b5e-e90a-4676-a5c8-0afa488daf1a · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.973955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.973955Z digest=sha256:48c23b6479d821827b8be5c0f395cb33c27b35f6df6ebff7f806e253c1241a17

Observation 3c43e290-8866-4bb6-b6bd-71e1d82d8cd9 · outbound

This paper cites Flexvar: Flexible visual autoregressive modeling without residual prediction.arXiv preprint arXiv:2502.20313, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flexvar: Flexible visual autoregressive modeling without residual prediction.arXiv preprint arXiv:2502.20313, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.976415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.976415Z digest=sha256:33cec972f65c5ffc2b20b4ab366445b7199e9b769bc321f9898d2cc6d474a1db

Observation c405f2ed-2f25-4181-8781-d7e6a222490c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.979380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.979380Z digest=sha256:3e2a8702cc2db03ee0632935f0ffe2a9834557656394e51c50af31557c071a9d

Observation 5cbb2112-205c-421a-a78b-d233b25cadff · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.981813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.981813Z digest=sha256:5aa156e280a16ce670e79df8a29260410f8b39ed0bd6e7bf66cdfe92219ff9d3

Observation 422361a2-1844-4d90-9333-4de7887b0b8f · outbound

This paper cites Kling ai: Next-generation ai creative studio.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Kling ai: Next-generation ai creative studio

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.722661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.984779Z digest=sha256:9ecbf421e2e4a3a3a933a070fb870b5239825ae8d65d298c14081af8100c05fd

Observation 94e46e44-a5ea-40bb-99f5-4746d64d10ce · outbound

This paper cites T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.714781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.987355Z digest=sha256:654404e4418660d4752ce1a58f5db63a14a0ab2954e194747f3058f662b1ce92

Observation 8ce437a4-5be9-4f9c-b721-a44e9b081b60 · outbound

This paper cites Video- mamba: State space model for efficient video understanding.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video- mamba: State space model for efficient video understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.706303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.989749Z digest=sha256:e3e4fc3d1a322b9e5f7513f720c8f09a2e3ee11423527c4626b9bdc2b12606dd

Observation 4863e008-1a78-4d8b-bbff-98e69c3a0b9f · outbound

This paper cites Mamba-nd: Selective state space modeling for multi-dimensional data.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-nd: Selective state space modeling for multi-dimensional data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.697737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:18.992237Z digest=sha256:737ea4b1e1e3b95f9af2190b692c43ed0709d9eed8e761c21f79d2c9e2ffd5b6

Observation 6f9e4cb6-9f41-4dc8-bb34-2620f93235b6 · outbound

This paper cites an unresolved cited work.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.995265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.995265Z digest=sha256:fcf64557193690484f7babcbbd0f690f42f9fcd47b63b74166f13b1725e10b58

Observation 7b53594f-2165-44be-81c2-4b4087fcd7d1 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.998136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.998136Z digest=sha256:f1ca2e5f4ee2bfa2faaf26607c7ba2d40353d269931902b400cb1b3d1c5117bc

Observation 023dacae-7917-461d-ae90-1b34abb3942a · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.001575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.001575Z digest=sha256:c20713077a74c14536b142421ea796be2b1cb94d5b3a5e13c92d835d181d8ebf

Observation 0a137ec2-b2a8-4e70-8aea-eb8bbe34de27 · outbound

This paper cites Ssm meets video diffusion models: Efficient video generation with structured state spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Ssm meets video diffusion models: Efficient video generation with structured state spaces

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.683804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.004269Z digest=sha256:8cb984504bd915c5024c13bee51522e662fe9dd725e1ff436e01a6db0385767d

Observation 30ace87d-e10e-4458-8963-d32eda9a9100 · outbound

This paper cites Scalable diffusion models with transformers.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.007012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.007012Z digest=sha256:f53c1109fc3c0d51f7257ad4b8f9034909ecf79b336fd48681abf16fc9208822

Observation 9baddb8e-b538-4417-a982-3ed0edcb66dc · outbound

This paper cites Open-sora plan.https://github.com/PKU-YuanGroup, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora plan.https://github.com/PKU-YuanGroup, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.667407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.009685Z digest=sha256:f16f9a93f7628f05a202fbe3a10bd7f0f0e94364e1298634839b91d68e1671ff

Observation e1e3ebb6-ce29-4691-909f-4ef3eba43ed9 · outbound

This paper cites Long-Context State-Space Video World Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Long-Context State-Space Video World Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.012307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.012307Z digest=sha256:52eaa81e3c897a9f2a28f2b61167ef7bacd7a491d96a9f95464148d5666fac91

Observation ecf29e05-1a33-40ec-9b37-7d0b6a5b09bb · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.015208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.015208Z digest=sha256:f38e4e9ebd6bd3d853d22d5685183c4550064c6c338bd3410ddd99fbce4b2c83

Observation 39fd6d6e-3bb8-41e0-a0bd-d9233b1fe3d7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Learning transferable visual models from natural language supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.018003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.018003Z digest=sha256:b6fd2fb9c122b53b4042cddb123b67c5d5436d3dd63e4f6055b37e16e1394253

Observation 80f53fd6-bffb-4f16-809e-235d7b0eef37 · outbound

This paper cites Runway gen-3 alpha.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Runway gen-3 alpha

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.652871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.020946Z digest=sha256:c3bf75f387c9f4e925c67fee9cb7babecf8939a6780f9f67ade840a6505ee893

Observation b61fb72a-947f-4612-96ff-5e04c7043793 · outbound

This paper cites Magi-1: Autoregressive video generation at scale, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Magi-1: Autoregressive video generation at scale, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.023647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.023647Z digest=sha256:f8fa60e50762cc7d2f4f2d4cb3418e6e2941fc3be7791e70ed30b343f3d6e353

Observation c1071c26-d5f1-4673-a5bb-b29185aafc57 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.026438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.026438Z digest=sha256:6a22d7f3891e4cd956833afe719a8751b41e2c365ef51ae3edd6245dd55c39a3

Observation 88867a5f-f8c2-4761-8f0b-6ce70adbd41a · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.029286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.029286Z digest=sha256:a88859ecb0da12d29a8fc2e85de9cd51e1615675172ea7214779943f2842c0bf

Observation 171c0843-9364-489c-96cc-fade8b2e79fe · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Diffusion Models Are Real-Time Game Engines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.031951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.031951Z digest=sha256:23103d0c78f90ea0a26552006aae60a646dd938139194c77a8e0dadb69a8612d

Observation e099cc7a-9241-4a4d-aac6-f382d5ea1d68 · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.035181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.035181Z digest=sha256:ed809713ad216fd717a60f42ccf4b3ab7dd254152e93d8ca9e16bb8ca740aebe

Observation 5d15457b-54ff-4f7f-a80a-bef5564c91a2 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.037975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.037975Z digest=sha256:fda5c60d9b9973bcc5b5cc479a1bc9307239b3b753c3e283143c47e8be992686

Observation 94382658-1a9c-421c-bc64-d13516b500fe · outbound

This paper cites Mamba-R: Vision Mamba ALSO Needs Registers.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-R: Vision Mamba ALSO Needs Registers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.040852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.040852Z digest=sha256:e0e2c69bfe100a900a25bdaf96900450d6e1ba53ad46a9e2e47afc70e625506a

Observation a8fe1376-17c3-4580-a606-0c671997fc7a · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.044277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.044277Z digest=sha256:b99d8b852b33e192797e01c9062e25cf749c001d62dd13717b1ea0758c025336

Observation 3f89ebbd-a47e-45c7-bfcc-a5ded8772ed8 · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.047077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.047077Z digest=sha256:fad2c4f742846894ae0be34800e39569a7783bd28173f5f019b08bf4f9a8eb67

Observation 3f63a038-4859-4885-8140-f52aeea1ed68 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.050020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.050020Z digest=sha256:50fff36dc1b5364031823fe88545529f5e838ba8d1f8e16c8c6f1617e4468c15

Observation f2d4b2fe-03ae-4eda-987a-aacd6d430732 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.052703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.052703Z digest=sha256:9312411f3441dff7d236aca18fae91fca5acf95f0f49d82a0fc6aa076c46681b

Observation 1af791d3-81fe-4bb7-9d33-0780994358d9 · outbound

This paper cites PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:21:19.210398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.055280Z digest=sha256:aff4d34d6cf2f77766ff6b0456a7bc4a3b9bd6d4d7a9235980cab8273f96deae

Observation d94db6d2-1b5f-4a17-b7fa-67133e896326 · outbound

This paper cites Perflow: Piecewise rectified flow as universal plug-and-play accelerator.Advances in Neural Information Processing Systems, 37:78630–78652, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Perflow: Piecewise rectified flow as universal plug-and-play accelerator.Advances in Neural Information Processing Systems, 37:78630–78652, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.620985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.058496Z digest=sha256:948d07d5b711691a529216d4c34a02119051b27999cf6cd8f9223da7d82c1dc4

Observation a7078d3f-bea1-4375-a03e-4bb3a1a766d7 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.061124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.061124Z digest=sha256:ef94adc5ccabb6a25c1979aec5a03717158d20dbf3ef12b9cf31cfa308226467

Observation 7481707c-e0bc-4de7-ba4d-7d075ff435f4 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.arXiv preprint arXiv:2412.07772, 2, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation From slow bidirectional to fast autoregressive video diffusion models.arXiv preprint arXiv:2412.07772, 2, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.063913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.063913Z digest=sha256:a87a76b77d92a6b4e695d8ce29d23ad525dbd6e6ab86f8afee9c963c831f9061

Observation e4ce0c0a-79f7-412e-9a5d-0d6c140088e9 · outbound

This paper cites Slca: Slow learner with classifier alignment for continual learning on a pre-trained model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Slca: Slow learner with classifier alignment for continual learning on a pre-trained model

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.611061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.066663Z digest=sha256:27a7e715b774107e0dc07802f023ad84eaf42d1c5b22b25d30de8aae18c51ff8

Observation 42402061-41fa-473c-a75e-9ddee9ec398c · outbound

This paper cites SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.069291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.069291Z digest=sha256:d7f2d05e648997f0355899bc967ddcb9346d785f9f406d9585e88df3ffdb7c9f

Observation 2619e31e-1d85-4800-a19d-cecf3ed49f9b · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora: Democratizing efficient video production for all, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.072026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.072026Z digest=sha256:e72677feca553b2cd9946bad807aa5d66129dc608880f7796cf79ac934a406f3

Observation 2b5bfff9-1850-44ec-8009-ef5e87941f61 · outbound

This paper cites downward first, then rightward.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation downward first, then rightward

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:21:19.596314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:21:19.076614Z digest=sha256:062e20da0cc5cc73bd8917d54b0a47277678ddda6c0953bdfe9c1fb3db87338d

Pith citing papers

Observation 1ddb49fa-c4ac-44ed-ba10-04f8fbe71211 · inbound

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving cites this paper.

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:19:42.932136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T19:19:42.748573Z digest=sha256:9efbdd5afe211ac54f707ff73d5047a995708f1d387be4ea4876d8eb21f725eb

Observation d5ca2794-f90d-48c7-bbc9-01ec5835e62d · inbound

Setting the Stage: Text-Driven Scene-Consistent Image Generation cites this paper.

Setting the Stage: Text-Driven Scene-Consistent Image Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.287078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:0918399d7c6a859b912214b3e9081b72eb95c23af6de48aa650fe35671b9a5a9

Observation b685d287-2d41-4a87-bf75-cbe6f457ba16 · inbound

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing cites this paper.

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T12:20:51.326210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:20:51.326210Z digest=sha256:5c1054e9491fe1ad5d4afe2ea29ff1f82e3bdb3860a91fa4986011c265029ea7

Observation c8844a78-64e8-4506-aee1-dfe9f0ec77da · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.767995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:acbf0cebc48d5400a4372ca385e5ab462841feaa7bf5edaca9f2c14d88b354fc

Observation 88582fc0-3acc-45a3-a988-110e9a377303 · inbound

MobileWan: Closing the Quality Gap for Mobile Video Diffusion cites this paper.

MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:44:59.795849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T14:37:46.957265Z digest=sha256:ff8885e2470ad4d5d77072cc2240707259c4f07c91d5bfb1fa8778f2b86589df

Observation a0e029a5-6f1e-44f4-881e-082c12d34802 · inbound

MobileWan: Closing the Quality Gap for Mobile Video Diffusion cites this paper.

MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:47.529772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:47.529772Z digest=sha256:e3dfdb6c5f8bf7d2fa09d0024959c27fa8a40db46a9d791a35ff315d3bc345fd