Pith. sign in

Paper Citation Record · LEDGER

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

As of 20 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 6 inbound Pith citation observations for arXiv:2506.10915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10915 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:21:19.076614Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:47.529772Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:44:59.794530Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7e0ff73-8461-4a52-aea9-cda721dab9a0 · outbound

This paper cites Pika art.https://pika.art, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pika art.https://pika.art, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.782774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.766104Z digest=sha256:ccac8e6f13451fe0601532b495be084a2831456b10d2c652e4d55ddda9eb2a30

Observation 02eda61e-eb01-498f-8d63-e897e9cd6f46 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.818385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.818385Z digest=sha256:f1ba5787e5f4be5d26ba33ffed7fa099eb9027fa075004895de485056a18a70a

Observation 99b9b93a-2590-49e7-8816-6e57030e277d · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2023.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flux.https://github.com/black-forest-labs/flux, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.769840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.821935Z digest=sha256:67d501d08a7899a762ad9caa7773a402e26c82eff4fe1fe83665a8b2e202a972

Observation 60d7f820-24d1-4853-af12-8bb281537e84 · outbound

This paper cites Video generation models as world simulators.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video generation models as world simulators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.923561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.923561Z digest=sha256:e62e10688810df092314039ea4aceea7a1dde9dfa5a95d4b5674f7548677c38a

Observation be5e0b4a-7155-47a8-a2dc-e306aed9dd50 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.927125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.927125Z digest=sha256:774147d3ff2c92c17843299db56c7214409a301ce72ed611f0005168ca225813

Observation ca55e251-000c-4ef7-aa51-3fa580c03782 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.934563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.934563Z digest=sha256:7676bd0edcabbf8c51f2a32f5b19606469d18fae76066ccb4e20c268de4c8c7f

Observation b2836212-e280-4e41-aa74-47852d59b0ec · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Goku: Flow Based Video Generative Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.937837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.937837Z digest=sha256:c6d0a68922d2c08393d991296bd98940819c4fa2f73ab19da8c88bb5fc97e1b0

Observation d86eef98-54df-4310-9284-7862710e3e4c · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.940770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.940770Z digest=sha256:3f4d2bf179c68b606b0cd8d7a6157cb8358af997c0763fb9b9df83c1b95ab8d3

Observation d2dbefb8-b6ce-4d60-add6-3c613861d226 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.943726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.943726Z digest=sha256:17aa64180e2710811ac29c52538b328d82063ac5fd1afc3079678deb992fef66

Observation 52f37071-c193-4085-8050-6f41ff3fbf83 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.946935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.946935Z digest=sha256:2fefac4ad81f57f4ce4315cc0171dc6b41769f0bc99deb40e71b1c04b09c8f9d

Observation e913141b-a757-4a09-a3fb-9a00293ffc7d · outbound

This paper cites Dimba: Transformer-Mamba Diffusion Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Dimba: Transformer-Mamba Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.950027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.950027Z digest=sha256:fdfb9f0891aadee7b444707f2731921ef1c4ae653c5a7150ef31b9008b0f77ee

Observation 47d6df64-ca99-4439-8328-8a0ffc12025e · outbound

This paper cites LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.953995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.953995Z digest=sha256:21b1ff2bfbfd20d5fd2eb9ede1430dff925514058222356fb7d2e78a378f6021

Observation ab791d2d-c81f-4b35-962c-1a12e938b4fa · outbound

This paper cites Matten: Video Generation with Mamba-Attention.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Matten: Video Generation with Mamba-Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.957450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.957450Z digest=sha256:b12e71a9ab5e26cf3417f3ccca5f06447851f33a6b4862dadfaf9f01d0a7992f

Observation d0462d30-c1b8-459d-be5f-85656edbcbc0 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.960505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.960505Z digest=sha256:344378876d66578fbeb911f6446e95682ab0329277d0a4f09344ed521015e0c1

Observation 27f80f9b-8858-4905-8597-1a94c3ee5496 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Efficiently Modeling Long Sequences with Structured State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.963248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.963248Z digest=sha256:e85ac020c1c851aec0e5786a674e588280cf02772d999d41b3088695f817b952

Observation 41e58600-93a5-434e-b8c8-c10c9be8c886 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.965955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.965955Z digest=sha256:7b9288f32ffc70fa78b673cbbd30b51774761ed09d46b24fe5649fbe99768a3e

Observation 21ac86b8-ebd5-4013-ba11-9b4cc914480f · outbound

This paper cites Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.968467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.968467Z digest=sha256:5ac720d615801453c6c661854b18e6432b1a55cb91be54616db8c4e7a4f95513

Observation 1989c988-afa6-4980-a0f8-6a9d576cd81c · outbound

This paper cites Zigma: A dit-style zigzag mamba diffusion model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Zigma: A dit-style zigzag mamba diffusion model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.735860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.970874Z digest=sha256:d3e674803c7fc3960362d55090c3e2d15c2a01cd99275ebba6480d1ec4e2e463

Observation 541e0b5e-e90a-4676-a5c8-0afa488daf1a · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.973955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.973955Z digest=sha256:71e80afd80da31509bde40495646741ff2fd9ba4d96ae94a304a2acda00cbbf4

Observation 3c43e290-8866-4bb6-b6bd-71e1d82d8cd9 · outbound

This paper cites Flexvar: Flexible visual autoregressive modeling without residual prediction.arXiv preprint arXiv:2502.20313, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flexvar: Flexible visual autoregressive modeling without residual prediction.arXiv preprint arXiv:2502.20313, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.976415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.976415Z digest=sha256:0052e459042908fa0accd731384341ca41b3a1c4e9a677044a31e2553a348fb5

Observation c405f2ed-2f25-4181-8781-d7e6a222490c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.979380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.979380Z digest=sha256:4eed0c4c90a66547bdcd509d8e0135b205772942bfb36d1a97b45ddd51cd500b

Observation 5cbb2112-205c-421a-a78b-d233b25cadff · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.981813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.981813Z digest=sha256:cc29422383b2c2da301d9b8cdc6d377c6a44bc4a14480c5bf9c959cbe82cb9f2

Observation 422361a2-1844-4d90-9333-4de7887b0b8f · outbound

This paper cites Kling ai: Next-generation ai creative studio.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Kling ai: Next-generation ai creative studio

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.722661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.984779Z digest=sha256:2960dd357ae22eef340a49a2764c2b2bea3afc9f6ee9c50dd8a3357a090e79e9

Observation 94e46e44-a5ea-40bb-99f5-4746d64d10ce · outbound

This paper cites T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.714781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.987355Z digest=sha256:5eb6f6c5d51910bea56ec63511a50199cb6dff461bb379c9ef4b58c50ee29125

Observation 8ce437a4-5be9-4f9c-b721-a44e9b081b60 · outbound

This paper cites Video- mamba: State space model for efficient video understanding.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video- mamba: State space model for efficient video understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.706303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.989749Z digest=sha256:034fd4cee5a77d77ff01b6f1270c5f1684a4e8849d09caee7fe0523714e3f2a5

Observation 4863e008-1a78-4d8b-bbff-98e69c3a0b9f · outbound

This paper cites Mamba-nd: Selective state space modeling for multi-dimensional data.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-nd: Selective state space modeling for multi-dimensional data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.697737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:18.992237Z digest=sha256:bcbd7e3c8b08f81d02a721ff49c538378a4c5c0a86937d7251b90d8ca0dee44a

Observation 6f9e4cb6-9f41-4dc8-bb34-2620f93235b6 · outbound

This paper cites an unresolved cited work.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.995265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.995265Z digest=sha256:90b760f2b4f8c38d58bbd0102751d896516298a5bb3919b2300fbf1be5ab3086

Observation 7b53594f-2165-44be-81c2-4b4087fcd7d1 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.998136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.998136Z digest=sha256:a295d99602af36f8837a156f17b1bad2e528c83c3777ce1c663bb7da08894158

Observation 023dacae-7917-461d-ae90-1b34abb3942a · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.001575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.001575Z digest=sha256:bf8523befdfea817a95e4a7ce8f2e40e5d5cbe0738277650f99c23491447b9b6

Observation 0a137ec2-b2a8-4e70-8aea-eb8bbe34de27 · outbound

This paper cites Ssm meets video diffusion models: Efficient video generation with structured state spaces.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Ssm meets video diffusion models: Efficient video generation with structured state spaces

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.683804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.004269Z digest=sha256:d2ec72082eaa026c81420af9aacb6127ce8a06353d33012db0b02e4c0b6d506e

Observation 30ace87d-e10e-4458-8963-d32eda9a9100 · outbound

This paper cites Scalable diffusion models with transformers.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.007012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.007012Z digest=sha256:77660059805822b345869bc2e9aee1dc30fb373cf93cd1886ec760796b2be52d

Observation 9baddb8e-b538-4417-a982-3ed0edcb66dc · outbound

This paper cites Open-sora plan.https://github.com/PKU-YuanGroup, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora plan.https://github.com/PKU-YuanGroup, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.667407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.009685Z digest=sha256:e63be6b4d1493a9390a61a62af89c7cc40babba32d29c9f77a8958aa8e725f9c

Observation e1e3ebb6-ce29-4691-909f-4ef3eba43ed9 · outbound

This paper cites Long-Context State-Space Video World Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Long-Context State-Space Video World Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.012307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.012307Z digest=sha256:e5547da08ebe15af2a432864e7f7b81b904f0343cf12843e90809a384489b5ec

Observation ecf29e05-1a33-40ec-9b37-7d0b6a5b09bb · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.015208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.015208Z digest=sha256:f300d92012427115585a255b4d082ee71ea36b1533429b5d95d230c43403cbf2

Observation 39fd6d6e-3bb8-41e0-a0bd-d9233b1fe3d7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Learning transferable visual models from natural language supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.018003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.018003Z digest=sha256:eeb189103d27975b6131c3e4309dd052f33fae9a3f0479261fac8204563bb827

Observation 80f53fd6-bffb-4f16-809e-235d7b0eef37 · outbound

This paper cites Runway gen-3 alpha.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Runway gen-3 alpha

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.652871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.020946Z digest=sha256:89cbd62dd2ecf7613067175431293f5df944bf04814d558f5bfc206bda4d32a1

Observation b61fb72a-947f-4612-96ff-5e04c7043793 · outbound

This paper cites Magi-1: Autoregressive video generation at scale, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Magi-1: Autoregressive video generation at scale, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.023647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.023647Z digest=sha256:396be79eafd621eb9b3edaa75ae58fdae533c2cccc6f50db62593f1fc7ec0874

Observation c1071c26-d5f1-4673-a5bb-b29185aafc57 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.026438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.026438Z digest=sha256:07db3ade90cfee299297d88cbff4251ddc60eee4a2fbad96a05f79d176b1c4c9

Observation 88867a5f-f8c2-4761-8f0b-6ce70adbd41a · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.029286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.029286Z digest=sha256:fb1a3dfe3982c7658e6fb1cb29ed20f2a7696ef59e642ed25f27ef8d4c19e1f0

Observation 171c0843-9364-489c-96cc-fade8b2e79fe · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Diffusion Models Are Real-Time Game Engines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.031951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.031951Z digest=sha256:35e3e9f01c99393ac398f96cd670780f1dda4c474c4d4a035b485203adf01352

Observation e099cc7a-9241-4a4d-aac6-f382d5ea1d68 · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.035181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.035181Z digest=sha256:71f42d9e72ed506921b0d2d4aad025eba1ae346d86daa7740b3970e4cbebd7d9

Observation 5d15457b-54ff-4f7f-a80a-bef5564c91a2 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.037975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.037975Z digest=sha256:de06c8f0dbd81a4c580eb909a89dbea3c5112721ce086d38e9eb126883081ab7

Observation 94382658-1a9c-421c-bc64-d13516b500fe · outbound

This paper cites Mamba-R: Vision Mamba ALSO Needs Registers.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-R: Vision Mamba ALSO Needs Registers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.040852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.040852Z digest=sha256:caffc57ad944bcb49fbe33166fdb2c8c69b3d7a447ea930e98339f900a413d04

Observation a8fe1376-17c3-4580-a606-0c671997fc7a · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.044277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.044277Z digest=sha256:b654acaff7c8b30a55d748e07320d94e7b287beb0a50629cc317384dbe819860

Observation 3f89ebbd-a47e-45c7-bfcc-a5ded8772ed8 · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.047077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.047077Z digest=sha256:7ef350a408926af365764ef698b93bc0f852367e41d4e295ef2d98c8ba9115b3

Observation 3f63a038-4859-4885-8140-f52aeea1ed68 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.050020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.050020Z digest=sha256:de317d3e88cba9bd28989ccbb8408423bbfec31b05ff45e5c8eb99e1643400eb

Observation f2d4b2fe-03ae-4eda-987a-aacd6d430732 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.052703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.052703Z digest=sha256:a8b0df7d2ff5992a0194628a33583a78ae366ea19da402de4c922c2b3ad595dd

Observation 1af791d3-81fe-4bb7-9d33-0780994358d9 · outbound

This paper cites PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:21:19.210398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.055280Z digest=sha256:38b0db8fb2c4e76e1c8c4e6c1cdaa64ee13876fe8e6f506ef842cd4e71cd7015

Observation d94db6d2-1b5f-4a17-b7fa-67133e896326 · outbound

This paper cites Perflow: Piecewise rectified flow as universal plug-and-play accelerator.Advances in Neural Information Processing Systems, 37:78630–78652, 2025.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Perflow: Piecewise rectified flow as universal plug-and-play accelerator.Advances in Neural Information Processing Systems, 37:78630–78652, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.620985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.058496Z digest=sha256:9f13150ea8392342eb8bd65f571ac5602ad6275d4d6924ef902f09ec7fd55551

Observation a7078d3f-bea1-4375-a03e-4bb3a1a766d7 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.061124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.061124Z digest=sha256:18a452bbd655efd691b82d4071a8cfc8b4bfa2b4b09b6ed92199c9e4fd906533

Observation 7481707c-e0bc-4de7-ba4d-7d075ff435f4 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.arXiv preprint arXiv:2412.07772, 2, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation From slow bidirectional to fast autoregressive video diffusion models.arXiv preprint arXiv:2412.07772, 2, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.063913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.063913Z digest=sha256:0e2cde1f75d47d888cfa0841d26869be5789f519da2af502a3283b57563e79f8

Observation e4ce0c0a-79f7-412e-9a5d-0d6c140088e9 · outbound

This paper cites Slca: Slow learner with classifier alignment for continual learning on a pre-trained model.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Slca: Slow learner with classifier alignment for continual learning on a pre-trained model

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:21:19.611061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.066663Z digest=sha256:102a21ca84f6696e03e26bae2b2c606614e04421eeb0e5810e03425e49443e32

Observation 42402061-41fa-473c-a75e-9ddee9ec398c · outbound

This paper cites SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.069291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.069291Z digest=sha256:6c1a0eb6eede81533eb1f08d84ff2ae209912e5e6e8db9fb3df9a89537b9af25

Observation 2619e31e-1d85-4800-a19d-cecf3ed49f9b · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora: Democratizing efficient video production for all, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.072026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.072026Z digest=sha256:768ddfc5d1c273297891a021ec51bfa60e72c0feba8da01ff6cfb2c03d86fca6

Observation 2b5bfff9-1850-44ec-8009-ef5e87941f61 · outbound

This paper cites downward first, then rightward.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation downward first, then rightward

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:21:19.596314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:21:19.076614Z digest=sha256:2f2fa1e2dea9703a0c1a3fa84a0af25902430fddf3a815b614679934034af9bf

Pith citing papers

Observation 1ddb49fa-c4ac-44ed-ba10-04f8fbe71211 · inbound

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving cites this paper.

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:19:42.932136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T19:19:42.748573Z digest=sha256:57518482b3cd146c036c44317f5d4002968ae9c60203a607cde66eaddd23d706

Observation d5ca2794-f90d-48c7-bbc9-01ec5835e62d · inbound

Setting the Stage: Text-Driven Scene-Consistent Image Generation cites this paper.

Setting the Stage: Text-Driven Scene-Consistent Image Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:44:17.287078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:970fda65fc1e7db867dce56f7e654d81630a25d91733999412f0f0f8e94b52f6

Observation b685d287-2d41-4a87-bf75-cbe6f457ba16 · inbound

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing cites this paper.

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T12:20:51.326210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:20:51.326210Z digest=sha256:92485b18362b01122e86553e81825f2a0fc6d45c54f32d2240cfecf2412be149

Observation c8844a78-64e8-4506-aee1-dfe9f0ec77da · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.767995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:107600032c80d0b97363c932109eb0831c40f954204050873d802e7f9b1fb976

Observation 88582fc0-3acc-45a3-a988-110e9a377303 · inbound

MobileWan: Closing the Quality Gap for Mobile Video Diffusion cites this paper.

MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:44:59.795849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T14:37:46.957265Z digest=sha256:9383b7d88c0d2714d5fe16dbd9d451ed3ec8fd2f9d9bed127eb329551d5b398a

Observation a0e029a5-6f1e-44f4-881e-082c12d34802 · inbound

MobileWan: Closing the Quality Gap for Mobile Video Diffusion cites this paper.

MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:47.529772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:47.529772Z digest=sha256:9223015c37591804484ecc8528d1a4f1643d32ef5d149066d1f7dfc37a618c3e