Pith. sign in

Paper Citation Record · LEDGER

Long-Context State-Space Video World Models

As of 11 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 13 inbound Pith citation observations for arXiv:2505.20171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20171 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:22.518133Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:21:19.012307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:57:31.679784Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7731cbc-afd5-4954-88d4-7e645be1f82c · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Long-Context State-Space Video World Models Diffusion for world modeling: Visual details matter in atari

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.094213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.094213Z digest=sha256:89781a41c8aae1b0b5e531204b3afa06bb665be2103996c4ae3c6fa3bda29132

Observation 964254a6-692d-4b64-af2a-653b8704fe62 · outbound

This paper cites Genesis: A universal and generative physics engine for robotics and beyond, 2024.

Long-Context State-Space Video World Models Genesis: A universal and generative physics engine for robotics and beyond, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.168532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.168532Z digest=sha256:2d26293e15ba311e60d284252ba2e15700a063db888339668d26ab9b71c55405

Observation 67468d09-0eca-4a29-9800-9db4a679f8d3 · outbound

This paper cites Layer Normalization.

Long-Context State-Space Video World Models Layer Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.313590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.313590Z digest=sha256:757ee0506c2e420f36a3fd9c7289119ca666bf1db0359d394c69f60fe88bfe50

Observation 847d4f8f-7ad3-4600-bcfa-9a0d6680fb12 · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Long-Context State-Space Video World Models Titans: Learning to Memorize at Test Time

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.404098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.404098Z digest=sha256:4f6f911a5be15e8a88d0c808c15de9d54b366f02799e512c4cc501dc60dc5df8

Observation 2736a171-0c7e-42d5-85d6-72382f192ac1 · outbound

This paper cites Decimamba: Exploring the length extrapolation potential of mamba.

Long-Context State-Space Video World Models Decimamba: Exploring the length extrapolation potential of mamba

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.550996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.550996Z digest=sha256:810bcafa17929f82990169f1e68e17696cc8b39742b641942cad011a2aa35d68

Observation 8277f684-58eb-43e6-b379-5c07c7dac116 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Long-Context State-Space Video World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.707554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.707554Z digest=sha256:a9d83af630fd9c3af4fec59ad2ffbc428b269984cfa00b45be8dbe8ddbd27317

Observation 48263f9d-3019-4c5f-823e-9e9d6cfe6075 · outbound

This paper cites Video generation models as world simulators.

Long-Context State-Space Video World Models Video generation models as world simulators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.815096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.815096Z digest=sha256:03a02bcde7f351f0e5f7dae0b88501bd163badb386d5e031f8e16a5e7e9e61d9

Observation 85323cea-7cd7-47e4-99fc-48aa6feff974 · outbound

This paper cites Genie: Generative Interactive Environments.

Long-Context State-Space Video World Models Genie: Generative Interactive Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.943294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.943294Z digest=sha256:85d8fad6529344088e97b2819b8ddfd951a5649638d3bf0ac2a241e75f3f5cef

Observation 519da0a0-627a-45d0-ac38-4110a4b10ce7 · outbound

This paper cites Gamegen-x: Interactive open-world game video generation.

Long-Context State-Space Video World Models Gamegen-x: Interactive open-world game video generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.045968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.045968Z digest=sha256:ac6778230b505e73a177517a2526115c58f34cd44d78be335e6b0c1e032f0b82

Observation a350047b-1105-4e18-9795-78ae5e989a81 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.

Long-Context State-Space Video World Models Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.137224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.137224Z digest=sha256:f319b3495c537fc3f428f4b83fabe0dc7642ddf8aa8f48f224306b910962871b

Observation b3d4e12c-c445-4b8a-8230-d2d6fe1d5ced · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

Long-Context State-Space Video World Models Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.267801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.267801Z digest=sha256:77c6e3757c6186bf8de254d3cdfaa118a821e9a32df28ba1853b6e42d25f55c8

Observation de8fa916-6d46-495e-8cd4-5c6bb12b91f5 · outbound

This paper cites Recurrent environment simulators.

Long-Context State-Space Video World Models Recurrent environment simulators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.415179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.415179Z digest=sha256:e891665ff2010b1427a502c995f97621ef8da1764e1f80dfe8257864362717ef

Observation ca41c59e-5d74-4159-9afc-93e34c6a85f3 · outbound

This paper cites Transformers are ssms: General- ized models and efficient algorithms through structured state space duality.

Long-Context State-Space Video World Models Transformers are ssms: General- ized models and efficient algorithms through structured state space duality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.542983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.542983Z digest=sha256:30e130af2a9f0ad05bfbb46519880c01ad9c065f283bfcc8700316ffc29623c1

Observation f3eaf4ae-9589-4be9-b561-8cddf27d34b3 · outbound

This paper cites Oasis: A universe in a transformer.

Long-Context State-Space Video World Models Oasis: A universe in a transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.660059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.660059Z digest=sha256:17f741cef81f829e4f6e9566efbff9411fc66bbe6f9acec1ac98cb6da17faeaa

Observation 222eed6a-2ac6-40aa-af1f-19e17cf49001 · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

Long-Context State-Space Video World Models Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.818750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.818750Z digest=sha256:6a4d70d10a5726b60938602de18e6866b7e79f394f9fc6c9e6bf8d1dd7c267ca

Observation a9deea28-4b67-434e-99ca-58a08032f7ab · outbound

This paper cites The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control.

Long-Context State-Space Video World Models The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.888852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.888852Z digest=sha256:1f4635e57fb010ebe51e3e65971c01a0e3d9b90a0d106549438b0e9dacb76a98

Observation a96975dc-4aeb-444e-9db7-10a4edf6f0c6 · outbound

This paper cites ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models.

Long-Context State-Space Video World Models ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:14.996741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:14.996741Z digest=sha256:752c7bc332e91e7c1aa79a488afde7eb7e1894d63d674697e94e852a2c1bd2d9

Observation f233286d-4875-406a-85c5-81bd97b25165 · outbound

This paper cites Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing.

Long-Context State-Space Video World Models Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:15.086030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:15.086030Z digest=sha256:ae314362345a665d588dc97fcabf42058501356d8558a4a211817e04d92ed616

Observation 875935de-12fb-40c6-b4f8-4ab3ed0569ff · outbound

This paper cites Matten: Video Generation with Mamba-Attention.

Long-Context State-Space Video World Models Matten: Video Generation with Mamba-Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:15.220599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:15.220599Z digest=sha256:7241a16705beb004a23f7d3e48b7dfcb51937bb56b1404a6a3d53643baf37415

Observation 96119e04-a918-49d7-a56e-1d141d6b69aa · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces.

Long-Context State-Space Video World Models Mamba: Linear-time sequence mod- eling with selective state spaces

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:30.837639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:15.307349Z digest=sha256:891e7fab18c4dc65e52add7ff7db22f50342058936c719b375e05491fc941e45

Observation f8fa790d-0392-4330-a37d-2fc7dd3c6305 · outbound

This paper cites Photorealistic video generation with diffusion models.

Long-Context State-Space Video World Models Photorealistic video generation with diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:30.604414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:15.432865Z digest=sha256:091447edd16f92be40758a7aecccfb096807e0ae7cba54ace1ff2396a345a649

Observation d075f22b-89c3-495d-a897-72d82bd78abe · outbound

This paper cites Recurrent world models facilitate policy evolution.

Long-Context State-Space Video World Models Recurrent world models facilitate policy evolution

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:30.404915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:15.585664Z digest=sha256:afce284ee106476b0005e4c449b4ddeae255cadfda1880acff5261ba867c408a

Observation 4be1f272-8c97-481d-be4f-a252e66f65c0 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Long-Context State-Space Video World Models LTX-Video: Realtime Video Latent Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:15.718498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:15.718498Z digest=sha256:fb5c3911f039a514f6281e3e1e4b21b77269ff995be80886d736d7b5aed79a0a

Observation 7b58a9b2-aa47-426b-9938-186a6051324f · outbound

This paper cites Dream to control: Learning behaviors by la- tent imagination.

Long-Context State-Space Video World Models Dream to control: Learning behaviors by la- tent imagination

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:30.282631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:15.840486Z digest=sha256:a388b7247ad5820cbcdf261bb2f868df1564941dc3c6200fb311a5e2efe4376d

Observation 8b281877-bc57-4655-8c14-5222c3132691 · outbound

This paper cites Pre-Trained Video Generative Models as World Simulators.

Long-Context State-Space Video World Models Pre-Trained Video Generative Models as World Simulators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:15.907426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:15.907426Z digest=sha256:2bd3caef9face83bc48eb5f210fe3b70b10f591475f324879db6b3cebe1b5209

Observation 2c3da88d-ba8f-4910-8aaa-847a5e295e7a · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Long-Context State-Space Video World Models Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.018949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.018949Z digest=sha256:cdb0fc63145669eca22276f30d719380d8782e4b7282752c215edd6db2f5d0a4

Observation a2b1bd05-859e-484f-8c0c-93152b1f8eb3 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Long-Context State-Space Video World Models Denoising diffu- sion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.160163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.160163Z digest=sha256:5409acf67f4484bd313525e9defee86fe55fa58f3818509b44f0ae661fef7726

Observation 989b68a3-13a7-4e59-af0a-29aea668730c · outbound

This paper cites Video dif- fusion models.

Long-Context State-Space Video World Models Video dif- fusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:30.023389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:16.257764Z digest=sha256:8bc7eeba20e208b0ad7d268b6ae597ea68d2099c64fc5e30d7dc45610a480291

Observation 7df3940e-ad96-4b7c-a8b5-06f557d3cd42 · outbound

This paper cites Long short-term memory.

Long-Context State-Space Video World Models Long short-term memory

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.331725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.331725Z digest=sha256:b94fb5a1841827d1befe2640d112e13bfaeffcf40a249215ca763e6f4fec3a05

Observation 229207fc-d208-4c04-ba2e-1ba4117056e1 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Long-Context State-Space Video World Models GAIA-1: A Generative World Model for Autonomous Driving

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.416898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.416898Z digest=sha256:a1151d94630ec620a9c313705e45ab001286c8c9d0476fa4def345db3b63210e

Observation 2ae7b6de-e6c6-4b77-bcfe-4164906c15f5 · outbound

This paper cites Acdit: Interpolating autoregressive con- ditional modeling and diffusion transformer.

Long-Context State-Space Video World Models Acdit: Interpolating autoregressive con- ditional modeling and diffusion transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.537343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.537343Z digest=sha256:ad2b599cd74f23f4788dfce4ffcf2033deaa0ef8d690ddd86b21f480813dfe3d

Observation aa26676e-58fe-4c3e-be36-3fe70df56d99 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Long-Context State-Space Video World Models Arbitrary style transfer in real-time with adaptive instance normalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:16.645420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:16.645420Z digest=sha256:04ceaa5b457d19a549fc4061d59ebb44559604f856c746a2cf3460802baeab20

Observation 4947d86f-3e5b-4ba3-8889-75aab7fd5cfd · outbound

This paper cites Multimodal unsupervised image-to-image translation.

Long-Context State-Space Video World Models Multimodal unsupervised image-to-image translation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:29.755571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:16.762878Z digest=sha256:f8d3aad3db104ec12f91075d4d3baa1ac8ecca8b12f6560a837bd1fe2721f78e

Observation 2e0614a8-c3b6-40ee-ad27-1ebfe8867088 · outbound

This paper cites Scope of va- lidity of psnr in image/video quality assessment.

Long-Context State-Space Video World Models Scope of va- lidity of psnr in image/video quality assessment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:29.514119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:16.894699Z digest=sha256:34921be248f23408de87c1a8840fc3e11b23cf4f503b46969e07ad092a9f209b

Observation 4920de98-84a5-498f-b4d9-fb662e427f41 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.

Long-Context State-Space Video World Models Pyramidal flow matching for efficient video generative modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:29.237917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:17.006540Z digest=sha256:324380f6647dcfa353c26bd74f7bd61df47c44aa28a8d8ac8798182124f30c9b

Observation 7ef85091-9189-43fe-a68d-28b5f17ec309 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Long-Context State-Space Video World Models How Far is Video Generation from World Model: A Physical Law Perspective

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:17.134702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:17.134702Z digest=sha256:f38d7f177349a9e98b38e502d5efa71a5c2eec541c6b6528f4cbff2cee65b77b

Observation e2b2dbc1-d9bf-43f3-86c4-5755133f6c0b · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Long-Context State-Space Video World Models Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.960909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:17.240504Z digest=sha256:28a95a97bd5dd0681b68273544776a0457c6b9a85867f3ec320008a2396034bd

Observation 54e1da3c-4805-491e-bd9d-60b3d74d0afb · outbound

This paper cites Auto-encoding varia- tional bayes.

Long-Context State-Space Video World Models Auto-encoding varia- tional bayes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.756823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:17.335865Z digest=sha256:fb8a63efcc09b30bbb8da85df960ec49b4dfe7f21f2c5098e6533c41eca95f68

Observation 5a3342d7-0d96-4021-8606-0150a1f12cfa · outbound

This paper cites Videopoet: A large language model for zero-shot video gen- eration.

Long-Context State-Space Video World Models Videopoet: A large language model for zero-shot video gen- eration

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.607509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:17.413062Z digest=sha256:b761f723fb55aa4e72af42a50117e840870db50c8509b3ebe8edf6bd93bd6b12

Observation d825f69f-92f4-48fd-9e3a-b70b20aab223 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Long-Context State-Space Video World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:17.521847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:17.521847Z digest=sha256:45e6c9c2c4367c9c044b78772df456b2712a0a0c441f67ecb0647280426a64c4

Observation 30acb6ad-bad2-4a6a-b9bb-f2209ae2aece · outbound

This paper cites Efficient spatially sparse inference for conditional gans and diffusion models.

Long-Context State-Space Video World Models Efficient spatially sparse inference for conditional gans and diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.379381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:17.605903Z digest=sha256:761cb3f7823129c54ede35773a004c7ec85eb450db7d0f767524c7aee8d7dab2

Observation 78d5b7f4-987c-41d8-8fff-0866a7dab333 · outbound

This paper cites CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up.

Long-Context State-Space Video World Models CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:17.722525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:17.722525Z digest=sha256:226155ccce5d5fbc79b534cbb22bd3449252062dfcf18109efbb02e0516fb12d

Observation 07f81108-c6cd-481c-b297-ffa44504f9ba · outbound

This paper cites LinFusion: 1 GPU, 1 Minute, 16K Image.

Long-Context State-Space Video World Models LinFusion: 1 GPU, 1 Minute, 16K Image

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:17.907050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:17.907050Z digest=sha256:e2f275f54a2047a51b2d41bcaa28068d4cfe784eea7340c0e3e8326925d25d11

Observation 8f078112-24ed-433d-a8ce-16f4e032b360 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Long-Context State-Space Video World Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:17.999337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:17.999337Z digest=sha256:ba6eecfbb8113b03a0de7b1d20df2920be09bb4655b219d19d7f262ba3bcee73

Observation 0db3d003-c977-4794-9ebb-df6f95fba207 · outbound

This paper cites Playable video gen- eration.

Long-Context State-Space Video World Models Playable video gen- eration

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.111106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:18.065670Z digest=sha256:d0be8dd08b3bf74d3554b65dc4291156a76f17d7a10497078265c5635f9b686b

Observation c4a193a0-a258-4691-a098-6bc60be42ac3 · outbound

This paper cites Trans- formers are sample-efficient world models.

Long-Context State-Space Video World Models Trans- formers are sample-efficient world models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:28.015404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:18.241996Z digest=sha256:ca92d4743e734f4a6f5f6a6889ec255bf8474a63906f257178e302ae92b2299a

Observation e7b5ca6c-05c2-4c64-945f-8f941d9f1212 · outbound

This paper cites Action-conditional video prediction us- ing deep networks in atari games.

Long-Context State-Space Video World Models Action-conditional video prediction us- ing deep networks in atari games

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:27.913432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:18.342889Z digest=sha256:7bcb6cfe406f71c7f380b94f5547a7c6e660b4ab1d079594303d8103823bd800

Observation 5e85fd6a-5274-429e-a918-0f5d03f88d01 · outbound

This paper cites SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces.

Long-Context State-Space Video World Models SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:18.416980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:18.416980Z digest=sha256:7599dfb5ac4dc1b50b7185c3e469c24baae39ac9d6de5148065256c8088d0042

Observation 652b46b9-da39-4458-a83b-e268701808e8 · outbound

This paper cites Genie 2: A large-scale foundation world model.

Long-Context State-Space Video World Models Genie 2: A large-scale foundation world model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:27.755835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:18.534361Z digest=sha256:a18dfa998183681a9029d8d8ee30d8e8ab1e2f0a7273c7a00718657bb9e2e094

Observation bdf3bfc0-fa1a-46ac-bf56-0277d90acc8d · outbound

This paper cites Evaluating Long-Term Memory in 3D Mazes.

Long-Context State-Space Video World Models Evaluating Long-Term Memory in 3D Mazes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:18.640529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:18.640529Z digest=sha256:071dea122771110f2dcf71bbbfc7ab4a039171354834ad9fa03f022b2b607f88

Observation 6ce10524-94fa-4064-a773-99840c330eec · outbound

This paper cites Scalable diffusion mod- els with transformers.

Long-Context State-Space Video World Models Scalable diffusion mod- els with transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:27.540600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:18.726109Z digest=sha256:6e739483022153d3e775dd950bc7c42a05f9fdfbf286f981f14c90d3e06ccd09

Observation 52a3b87a-82de-4dc9-84ff-ed35d2ef92f3 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Long-Context State-Space Video World Models Movie Gen: A Cast of Media Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:18.814503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:18.814503Z digest=sha256:176fa687c7500f76715c0d48aa5878152259d513fb14b9b822ddac640c60f991

Observation 8a254fc4-8ee8-4ec4-9b55-a4f97642db41 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Long-Context State-Space Video World Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:18.919699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:18.919699Z digest=sha256:0ce92e69cae7e3ebc95cc60e4703c3f3f5a59c652733abe9c1d8ef116256722f

Observation 18fd7d6d-d3fd-413a-840d-280a2be88e76 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

Long-Context State-Space Video World Models Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.068672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.068672Z digest=sha256:e83265337c3a0e761b12f8c074edf55828b76d6cfcffa78bf31d1eccc980a65a

Observation c8c3a4a9-1ccb-4dcb-a202-43f934d3c21e · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Long-Context State-Space Video World Models High-resolution image syn- thesis with latent diffusion models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.141580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.141580Z digest=sha256:35374b47a4052ca00201caf0a11dd87b76f58e45795eae5c6edcf760dc6b9429

Observation c3122658-0588-4170-81dc-8a1cc94d59be · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Long-Context State-Space Video World Models Linear transformers are secretly fast weight programmers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.248197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.248197Z digest=sha256:6776f38698d30469e5eae8e4e592c4fb2cf30994fa3cdfde10a0f36a013c6353

Observation 17bb820e-575e-48fa-8378-cc3b6e1f0a1c · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Long-Context State-Space Video World Models Make-a-video: Text-to-video generation without text-video data

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:27.223712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:19.357819Z digest=sha256:0e16a86ffb7233cb6ece74101c7d0fdd5a0bf0da3bc66d39e7497b5ef3ca25f5

Observation f119f5f5-fe32-41c8-b5b0-ab2101192b55 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Long-Context State-Space Video World Models Deep unsupervised learning using nonequilibrium thermodynamics

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.439809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.439809Z digest=sha256:13fafba0ea33044bb9c4f83e1387b1a63c94d5b2cf38cfa6d111c87b80fe1c82

Observation b458a00d-8161-4a70-9e19-c6e5044fe715 · outbound

This paper cites History-Guided Video Diffusion.

Long-Context State-Space Video World Models History-Guided Video Diffusion

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.511480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.511480Z digest=sha256:246972f4760937b64696e3a39150b79613417fa962d5328e179b6533ffbae006

Observation 5053d318-0ec0-4cde-b52e-5c080c08c12b · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

Long-Context State-Space Video World Models Score-based generative modeling through stochastic differential equa- tions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.619752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.619752Z digest=sha256:7212dd60f797ec10a67dad29f4378cd934ca9a437f4c5ab3f509537e4bdc3ce9

Observation 4ce5bba6-1ad2-4c11-ba80-3c233cd72ce5 · outbound

This paper cites Me playing a few minutes of ai minecraft.

Long-Context State-Space Video World Models Me playing a few minutes of ai minecraft

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:26.952215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:19.726882Z digest=sha256:3339e92c2a06eeb0506d2c0ee26fd66ec810de7cf3d96ee59f7d36f488750219

Observation d07da5bd-7601-433b-993f-05070d5f1fe8 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Long-Context State-Space Video World Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.858698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.858698Z digest=sha256:2c6b8602016aa3c42f2a5521acaebc2124342354d85c3c2e83bd106b74c214c2

Observation 70d73b0f-5669-4fb1-ac8c-e3dada902d9c · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Long-Context State-Space Video World Models Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:19.946773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:19.946773Z digest=sha256:5b93471baed0bdd3bdb0a98df1f6a053413561af33913edcc01485d19ac0081d

Observation e6424556-1131-4bd9-a5d7-be9a11f3d745 · outbound

This paper cites DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis.

Long-Context State-Space Video World Models DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.027371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.027371Z digest=sha256:19ae5000c852c9a4baf23a00bd74d8ceccc66241824d5cb849c80b715d62e62c

Observation 9312dbd3-da2f-4a60-b2a0-2764361bf3d3 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Long-Context State-Space Video World Models Diffusion Models Are Real-Time Game Engines

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.157040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.157040Z digest=sha256:c2ef38534b110eb776c906391f29735171b676b5940a8a4da447fbb064105d55

Observation e28f9ca7-2a03-4b96-8781-7015805f778c · outbound

This paper cites Attention is all you need.

Long-Context State-Space Video World Models Attention is all you need

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.254830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.254830Z digest=sha256:32986a5eb7507d85810b101679d260b6629262979bd819a4c5dd03f1b5c9a5d8

Observation 41ee4b61-a540-401a-8fc6-08320cdc7712 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Long-Context State-Space Video World Models Phenaki: Variable length video generation from open domain textual descriptions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:26.737810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:20.366337Z digest=sha256:b185d39b3d993546df8e390687e8b26e3bb9511aebe7df9d3001590132c12242

Observation 0b237622-2fa2-45f5-8386-99bd91f8cf11 · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

Long-Context State-Space Video World Models LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.460990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.460990Z digest=sha256:da2ba5f7ef2e9b65a4c1c055a4121140a63e875cfc760cbadcbb09a3632580f1

Observation 909be15d-779c-4c8f-9a89-a3e915be256c · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

Long-Context State-Space Video World Models WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.562826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.562826Z digest=sha256:4be15d8f94ae986d0989459dec5f4c39fe28d597123684d64fc552c3548fad66

Observation 234ef67b-9fed-4f69-b73b-0b10d0e4e238 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

Long-Context State-Space Video World Models Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.688743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.688743Z digest=sha256:17545bfc9480b49832f32f3459c3cbec0bb49322b357a912ccf643f2f4f5431b

Observation 1388b870-4fd4-46a3-a927-bd729ff9cf2e · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Long-Context State-Space Video World Models Image quality assessment: from error visibility to structural similarity

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:26.533148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:20.771057Z digest=sha256:ab2b15356b853c386c5d329efc390e55564b061d3a9b4be6618c61415ba232d6

Observation 11c6c57c-bfa3-4e60-9e71-1e81bb80c769 · outbound

This paper cites Art-v: Auto-regressive text-to- video generation with diffusion models.

Long-Context State-Space Video World Models Art-v: Auto-regressive text-to- video generation with diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:26.300393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:20.839420Z digest=sha256:32f8a63e4c5b70cafac459d46cecd67226a82f5728d992a50db495b6827bd5de

Observation 309a7e9d-e34a-4216-a77b-643760c35dbd · outbound

This paper cites Daydreamer: World models for physical robot learning.

Long-Context State-Space Video World Models Daydreamer: World models for physical robot learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:26.119563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:20.900150Z digest=sha256:cb883d008b7132bea11f5c4282a3d6ed570e5bdab01a7b1c3584830615f10991

Observation aa128de2-289c-4198-b136-a9ef32aed6f5 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

Long-Context State-Space Video World Models Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.983153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.983153Z digest=sha256:608fba6549e7343893816b7bc8117fa928f4d125194055bc49eef7e86bf8cc33

Observation 7304587a-a298-4ce8-ae4a-3ee3e664744f · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Long-Context State-Space Video World Models SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.075820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.075820Z digest=sha256:89bf8325e06447531c75e587484c0e784cd4f28db6ffa2a0c9a20d98f676b705

Observation 388b8e3c-c83f-4e93-9827-0de7c9324025 · outbound

This paper cites Diffu- sion models without attention.

Long-Context State-Space Video World Models Diffu- sion models without attention

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:25.917127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:21.192966Z digest=sha256:9fe3721a5fc30cc2dc9a1257883062048029d39150f403ee69055371e3d4d024

Observation 6d0575d1-3475-47b8-a4cb-7a7af02e807a · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Long-Context State-Space Video World Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.272106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.272106Z digest=sha256:735fab652f432d6c1d16e61774c9c5c0b7338483e28623ebd89f0761d0aeefa1

Observation 7a665ead-fc74-4896-a9c8-b819fc1adb39 · outbound

This paper cites an unresolved cited work.

Long-Context State-Space Video World Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:03:25.757592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:21.364568Z digest=sha256:228a8b7ec3089b5085420a394eab4a53e7ce70888eebc03e174821da629abb7e

Observation 7e21047d-a472-4e77-ac10-d86132592f56 · outbound

This paper cites Learning Interactive Real-World Simulators.

Long-Context State-Space Video World Models Learning Interactive Real-World Simulators

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.470654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.470654Z digest=sha256:04af363922f3ebf663e26abf798775a0bc5610a443a5295873c5a9feae177556

Observation 56b7215f-d2ce-4be3-8b47-9003c2e669bb · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Long-Context State-Space Video World Models Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.536208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.536208Z digest=sha256:759fd8cda2b93b7da251e8b765e2c84a701e7e18dae3a6366102f3e0ace30c72

Observation 740bc463-2b72-4d8b-85fd-3318b8d55a5b · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Long-Context State-Space Video World Models Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.630393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.630393Z digest=sha256:821bdcdb8db2fcd0282b83705d76467e700f4888d073249487443f4f291cf267

Observation d7ec5cd7-723c-45c3-a4c9-b42de1df83c1 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

Long-Context State-Space Video World Models Parallelizing linear transformers with the delta rule over sequence length

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:25.603899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:21.734094Z digest=sha256:4e2baf0765fb8ed79bfeb50f2553660406514ea05f3c0540a08fdd017dc4a153

Observation 6f2480ad-fac2-4c7b-a1eb-186157e1c190 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Long-Context State-Space Video World Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:21.815035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:21.815035Z digest=sha256:44e90cfc380f9a981ada098092ba0536c749b8a48593eb1be911c44f24f1f4d3

Observation bf691532-5f12-4e48-8d8c-b90cb104d130 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

Long-Context State-Space Video World Models Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:25.382778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:21.918729Z digest=sha256:2a6376a7fd3e1d0eb0b899229929f58a2c5cc4be74fad85d0f7018609b89bebe

Observation 91970bd0-59f5-4365-8dd8-ba581a3249c4 · outbound

This paper cites Longmamba: Enhancing mamba’s long-context capabilities via training-free receptive field en- largement.

Long-Context State-Space Video World Models Longmamba: Enhancing mamba’s long-context capabilities via training-free receptive field en- largement

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:25.164508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:22.027914Z digest=sha256:59d4c694e43c1a7c9e4df7a67c4b33c3e5a0f0064ea7eadb505594dbaa9af860

Observation 6c7818fc-e661-48f3-b569-e971e4c66ad1 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion mod- els.

Long-Context State-Space Video World Models From slow bidirectional to fast autoregressive video diffusion mod- els

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:25.013960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:22.125610Z digest=sha256:ace4f7e80c67f2ad2ce0933403466284f64a59e208676aeda566da73b0dc8b58

Observation ca417230-ed20-4d09-b209-d797aa74d254 · outbound

This paper cites Gamefactory: Creating new games with gen- erative interactive videos.

Long-Context State-Space Video World Models Gamefactory: Creating new games with gen- erative interactive videos

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:22.227097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:22.227097Z digest=sha256:af92b317f5592ea60ac9b4ab40dc13dae0eedf2723eceb33fb137f52b971d858

Observation b74f5836-1642-4e11-9722-30fdf273a8d8 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Long-Context State-Space Video World Models The unreasonable effectiveness of deep features as a perceptual metric

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:22.353236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:22.353236Z digest=sha256:e818ed218d005f438988ce6635d951f2785be9cfde8ee641ce656edb1daa35aa

Observation bbfcc725-c9e3-4758-b089-447b952957d2 · outbound

This paper cites Extdm: Distribution extrapolation diffu- sion model for video prediction.

Long-Context State-Space Video World Models Extdm: Distribution extrapolation diffu- sion model for video prediction

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:24.797890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:03:22.453245Z digest=sha256:da2a3a93fb270ae8ff0e3bd351c6c01567c202cf6f5e771a17bfab03b0abde97

Observation d4934b74-e8b9-4767-9c49-d9d3fc886581 · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

Long-Context State-Space Video World Models Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:22.518133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:22.518133Z digest=sha256:bce32435e46cde04fd378151ad797549235320c7a35ea3dea5eedbc505e87f9a

Pith citing papers

Observation 1773a11a-af1e-4a1a-bcd0-d78ee8101786 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion Long-Context State-Space Video World Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:36:53.152030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:d1bb03d144afda38ccfb8273ed1900324548c65edc8e6f3ebdcfcd641c7e3fa6

Observation e1e3ebb6-ce29-4691-909f-4ef3eba43ed9 · inbound

M4V: Multimodal Mamba for Efficient Text-to-Video Generation cites this paper.

M4V: Multimodal Mamba for Efficient Text-to-Video Generation Long-Context State-Space Video World Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:19.012307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:19.012307Z digest=sha256:2803c82d796d00a382046d71fe44375c4d168f0166d6a283e783ab3bb573070b

Observation 9108f35e-89c0-4798-9775-7ec0528065f9 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Long-Context State-Space Video World Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:06.711779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:67c3034ac061eb097d2ba06a99c5a2ef71d11ddc9f32f385721380926ea15dd0

Observation 3527f628-1351-4165-b6a4-9859dfeef824 · inbound

Matrix-game 2.0: An open-source real-time and streaming interactive world model cites this paper.

Matrix-game 2.0: An open-source real-time and streaming interactive world model Long-Context State-Space Video World Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:36:53.375385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T22:36:31.044743Z digest=sha256:45c22b86c1c52c7721ca9b973a04445f79908ed508ba20f0cf9c635947cb13f6

Observation ca1d08b3-b425-42fa-a9f8-839af546d97e · inbound

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling cites this paper.

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling Long-Context State-Space Video World Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T15:46:10.265242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:46:10.265242Z digest=sha256:26dd58ac189e1cfb04e0c93583146168ca7c3cf566ebcd5e04e59085fa1b862c

Observation 7fc7d00f-f2e5-4cdf-be23-22e04d068833 · inbound

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration cites this paper.

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Long-Context State-Space Video World Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:18:18.330378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:15:54.413960Z digest=sha256:df01e7e1d59bdedbe4f4ac6fbd3732df96c638a64597cc782b5c3299082699f6

Observation 339c9c00-4d60-49ab-9ac4-417224d0d756 · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation Long-Context State-Space Video World Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:08:13.663975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:bd3b2f3362da196ca662037db2d2b2d11f9319b7eb4f802978fbba1075fac93b

Observation 2e007854-4e2f-4b2f-bf2c-91a5db6d3f41 · inbound

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications cites this paper.

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications Long-Context State-Space Video World Models

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.753547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:36:23.776293Z digest=sha256:d80d32a21263c0acbefc18f2658c7260a5a693d3bc32187e890e3bc2b3ec128a

Observation 0f336968-a046-4857-9dd6-f4e569695d36 · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models Long-Context State-Space Video World Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.774285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:8d9b901b631614dcc2c507f2e315d6f9c08e904fdf77ac4714c84ab27c6737e6

Observation 8e6e4ec2-3108-4eb0-859a-025bfb648964 · inbound

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory cites this paper.

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory Long-Context State-Space Video World Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.997356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:58:47.490871Z digest=sha256:e2e31c9e4c5d16e18ca12d01746c712d8375b29df0891908aefa08bfb678d5e4

Observation 6155cf46-a12b-47ed-9401-d555d081b73f · inbound

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory cites this paper.

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory Long-Context State-Space Video World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T17:14:54.438474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:14:54.438474Z digest=sha256:5358b4abb8ba5c255b16632640a4573c8f18b89b8b6fd2d241b9d335d486d638

Observation aa774030-1094-409a-b8ea-c2344f10c2e0 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Long-Context State-Space Video World Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.872252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:c0ff5d3e33440fc6034f8bb87b545b42f9aacb809e1f962514733aa29da212e4

Observation a6773206-7a53-4f72-b052-7e934bd17cf1 · inbound

Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models cites this paper.

Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models Long-Context State-Space Video World Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:57:31.681222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T18:48:29.572602Z digest=sha256:7376313ba93c40c2967c72e3a8b2583e5bad92496cdb433525ba097994e76b7d