Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:21:19.076614Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 6 inbound Pith citation observations for arXiv:2506.10915.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:21:19.076614Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:47.529772Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T14:44:59.794530Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f7e0ff73-8461-4a52-aea9-cda721dab9a0 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pika art.https://pika.art, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02eda61e-eb01-498f-8d63-e897e9cd6f46 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b9b93a-2590-49e7-8816-6e57030e277d · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flux.https://github.com/black-forest-labs/flux, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 60d7f820-24d1-4853-af12-8bb281537e84 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video generation models as world simulators
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5e0b4a-7155-47a8-a2dc-e306aed9dd50 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca55e251-000c-4ef7-aa51-3fa580c03782 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2836212-e280-4e41-aa74-47852d59b0ec · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Goku: Flow Based Video Generative Foundation Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86eef98-54df-4310-9284-7862710e3e4c · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2dbefb8-b6ce-4d60-add6-3c613861d226 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scaling rectified flow transform- ers for high-resolution image synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f37071-c193-4085-8050-6f41ff3fbf83 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e913141b-a757-4a09-a3fb-9a00293ffc7d · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Dimba: Transformer-Mamba Diffusion Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d6df64-ca99-4439-8328-8a0ffc12025e · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab791d2d-c81f-4b35-962c-1a12e938b4fa · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Matten: Video Generation with Mamba-Attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0462d30-c1b8-459d-be5f-85656edbcbc0 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f80f9b-8858-4905-8597-1a94c3ee5496 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Efficiently Modeling Long Sequences with Structured State Spaces
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e58600-93a5-434e-b8c8-c10c9be8c886 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ac86b8-ebd5-4013-ba11-9b4cc914480f · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1989c988-afa6-4980-a0f8-6a9d576cd81c · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Zigma: A dit-style zigzag mamba diffusion model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 541e0b5e-e90a-4676-a5c8-0afa488daf1a · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Vbench: Comprehensive benchmark suite for video generative models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c43e290-8866-4bb6-b6bd-71e1d82d8cd9 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Flexvar: Flexible visual autoregressive modeling without residual prediction.arXiv preprint arXiv:2502.20313, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c405f2ed-2f25-4181-8781-d7e6a222490c · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbb2112-205c-421a-a78b-d233b25cadff · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 422361a2-1844-4d90-9333-4de7887b0b8f · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Kling ai: Next-generation ai creative studio
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94e46e44-a5ea-40bb-99f5-4746d64d10ce · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation T2v-turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ce437a4-5be9-4f9c-b721-a44e9b081b60 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Video- mamba: State space model for efficient video understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4863e008-1a78-4d8b-bbff-98e69c3a0b9f · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-nd: Selective state space modeling for multi-dimensional data
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6f9e4cb6-9f41-4dc8-bb34-2620f93235b6 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b53594f-2165-44be-81c2-4b4087fcd7d1 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023dacae-7917-461d-ae90-1b34abb3942a · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a137ec2-b2a8-4e70-8aea-eb8bbe34de27 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Ssm meets video diffusion models: Efficient video generation with structured state spaces
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 30ace87d-e10e-4458-8963-d32eda9a9100 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Scalable diffusion models with transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9baddb8e-b538-4417-a982-3ed0edcb66dc · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora plan.https://github.com/PKU-YuanGroup, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1e3ebb6-ce29-4691-909f-4ef3eba43ed9 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Long-Context State-Space Video World Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf29e05-1a33-40ec-9b37-7d0b6a5b09bb · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39fd6d6e-3bb8-41e0-a0bd-d9233b1fe3d7 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Learning transferable visual models from natural language supervision
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f53fd6-bffb-4f16-809e-235d7b0eef37 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Runway gen-3 alpha
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b61fb72a-947f-4612-96ff-5e04c7043793 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Magi-1: Autoregressive video generation at scale, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1071c26-d5f1-4673-a5bb-b29185aafc57 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems, 35:25278–25294, 2022
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88867a5f-f8c2-4761-8f0b-6ce70adbd41a · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 171c0843-9364-489c-96cc-fade8b2e79fe · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Diffusion Models Are Real-Time Game Engines
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e099cc7a-9241-4a4d-aac6-f382d5ea1d68 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Attention is all you need.Advances in Neural Information Processing Systems, 2017
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d15457b-54ff-4f7f-a80a-bef5564c91a2 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94382658-1a9c-421c-bc64-d13516b500fe · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Mamba-R: Vision Mamba ALSO Needs Registers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fe1376-17c3-4580-a606-0c671997fc7a · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f89ebbd-a47e-45c7-bfcc-a5ded8772ed8 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation The Mamba in the Llama: Distilling and Accelerating Hybrid Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f63a038-4859-4885-8140-f52aeea1ed68 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d4b2fe-03ae-4eda-987a-aacd6d430732 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Easyanimate: A high-performance long video generation method based on transformer architecture.arXiv preprint arXiv:2405.18991, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af791d3-81fe-4bb7-9d33-0780994358d9 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d94db6d2-1b5f-4a17-b7fa-67133e896326 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Perflow: Piecewise rectified flow as universal plug-and-play accelerator.Advances in Neural Information Processing Systems, 37:78630–78652, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a7078d3f-bea1-4375-a03e-4bb3a1a766d7 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7481707c-e0bc-4de7-ba4d-7d075ff435f4 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation From slow bidirectional to fast autoregressive video diffusion models.arXiv preprint arXiv:2412.07772, 2, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ce0c0a-79f7-412e-9a5d-0d6c140088e9 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Slca: Slow learner with classifier alignment for continual learning on a pre-trained model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 42402061-41fa-473c-a75e-9ddee9ec398c · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2619e31e-1d85-4800-a19d-cecf3ed49f9b · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation Open-sora: Democratizing efficient video production for all, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5bfff9-1850-44ec-8009-ef5e87941f61 · outbound
M4V: Multimodal Mamba for Efficient Text-to-Video Generation downward first, then rightward
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ddb49fa-c4ac-44ed-ba10-04f8fbe71211 · inbound
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5ca2794-f90d-48c7-bbc9-01ec5835e62d · inbound
Setting the Stage: Text-Driven Scene-Consistent Image Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b685d287-2d41-4a87-bf75-cbe6f457ba16 · inbound
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8844a78-64e8-4506-aee1-dfe9f0ec77da · inbound
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88582fc0-3acc-45a3-a988-110e9a377303 · inbound
MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a0e029a5-6f1e-44f4-881e-082c12d34802 · inbound
MobileWan: Closing the Quality Gap for Mobile Video Diffusion M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.