Pith. sign in

Paper Citation Record · LEDGER

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

As of 18 August 2026, this Paper Citation Record lists 100 of 116 outbound references and 100 inbound Pith citation observations for arXiv:2311.15127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.15127 v1

Coverage vector

measured 100 of 116 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T22:58:51.792047Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 100 of 1074 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:41:59.389473Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 116 outbound references displayed

  • verified exact32
  • verified fuzzy60
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

67
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5da6655c-ed78-40ef-bafd-3a4f839b1168 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.156087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b12deeaee9073fa6ef3f85bd14a49f2cefa8f3bb7e68812dfaa727d5d1cd680f

Observation 144bcf58-117e-4840-8e26-ea1f0d350bab · outbound

This paper cites Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.351279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4543fced0a9e9b41507cc484baebd63f84cfabffe73e5fb5d893076fa7d6fd3d

Observation 1369ee33-e07b-4808-a922-19ef6454c28c · outbound

This paper cites A general language assistant as a laboratory for alignment.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets A general language assistant as a laboratory for alignment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.354002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b3b88676e7a63e243d7f38a1cf0c53dc7c284356ddb118a2f4dc02f7c7ab67eb

Observation 93cc8df8-95fc-4c9f-be29-a1da5d5b8662 · outbound

This paper cites Campbell, and Sergey Levine.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Campbell, and Sergey Levine

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.356405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:33f483408993da00f4d983727bb509ae651b9b427c40167e22e98c9910ed93fb

Observation 933ef10c-4406-4898-814a-9df65483ee6b · outbound

This paper cites Character region awareness for text de- tection.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Character region awareness for text de- tection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.360851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:ba0e453f33d187ae1d5345709c7c28e8066445a3811f6571f33795be2dd7dff1

Observation db7d9742-6bab-457f-821b-97f97f659745 · outbound

This paper cites Training a helpful and harm- less assistant with reinforcement learning from human feed- back.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Training a helpful and harm- less assistant with reinforcement learning from human feed- back

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.367993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:586cde3877594747e3ff6921d0fc6747c145a76d5a22092263db4fa369f59bb6

Observation 22aa8867-91ab-44a3-a949-85e04be5d2bb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.372008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:94109fdfc73ed7c9585d6847d531b979f3ba2ef489dda11e6465e64f15ddb1e2

Observation 50404684-73cb-43f0-aa23-7cf3fbb80143 · outbound

This paper cites ipoke: Poking a still image for con- trolled stochastic video synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets ipoke: Poking a still image for con- trolled stochastic video synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.374897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f8fafc27b290bc06d807b09741dde5d67866a244da8a85f344053a7a825b9b62

Observation 9911c2c8-7ca3-40c2-a592-bbe9e7d0d173 · outbound

This paper cites Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.140275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f391a6c39ff67d0b77d60b6b819d57d165c0704014cb39fa6846aeae312975fb

Observation 560ec6c5-ca90-4822-bdfa-d8858b739e0e · outbound

This paper cites Generating long videos of dynamic scenes.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generating long videos of dynamic scenes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.380162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b3e758f4f4279c2e420284688bded41db19ea93482f0613fe3cbbcbfbc3bca40

Observation 84289ebc-a5e0-4436-98b1-657f7e6aac5a · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Quo vadis, action recognition? a new model and the kinetics dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.383092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a3e21bad3c8c5f143a862bc5129593eddc692e1b312420a39d6798a7e4f44aea

Observation 3df0282d-bca6-472d-93e9-a5c7ff611852 · outbound

This paper cites Im- proved conditional vrnns for video prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Im- proved conditional vrnns for video prediction

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.389104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c44fcdd04c74a39aeb5cb0a046c487c591f7a4515dc43ac901180c3f073d0690

Observation 3a9c73d5-ff6c-42ac-99ee-0e9da43b4573 · outbound

This paper cites Emu: Enhancing image generation models using photogenic nee- dles in a haystack.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Emu: Enhancing image generation models using photogenic nee- dles in a haystack

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.395124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f2c2f45ba2ca05749a7da48118c3d1b7c8c9ee47495975c421feea92674678c2

Observation d0f5b2f2-63ce-40e6-943e-a29cb2343dad · outbound

This paper cites Objaverse-XL: A Universe of 10M+ 3D Objects.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Objaverse-XL: A Universe of 10M+ 3D Objects

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:02:11.861391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7e506bc1e1424ba50572bff2d67fec32b3e9b5827c95072b4ee4ce4d693df7b1

Observation bd7479d9-f02e-4934-9fdd-5e7b8e090b5d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Objaverse: A universe of annotated 3d objects

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.398112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3478b42c286dc3c54a17f723825a0fd97b417ebce886f9468f5f05176d1b1d85

Observation 5a104dc8-f20a-4b87-be90-f5352dab01a2 · outbound

This paper cites Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.401071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:493c0f4e48acd46e7bcd948f062dcf284b8228a0a375cd0f283c3a3f5cf39af4

Observation 320fb33a-6144-4d89-9f0b-560514fdc576 · outbound

This paper cites Stochastic video genera- tion with a learned prior.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic video genera- tion with a learned prior

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.404300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:12be6ccdacba3f6652fa323c796a0a986d0e3b6370d760c0e4ab328763403f11

Observation 6b2076e6-a43b-4069-99c3-f9aa92e5dbba · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Diffusion Models Beat GANs on Image Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:16:28.664120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:87bd6d311f63539849df4114de201497c9dce6f90a37a66c4f74f320c81c5eb6

Observation 020a0ee9-fbcb-44c6-af09-1542b02f5622 · outbound

This paper cites Derpanis, and Bj¨orn Om- mer.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Derpanis, and Bj¨orn Om- mer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.406918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4493cc9f4b520c31d310b5a1e0a4fb46868f91f0c19ad0bb9b909229e0123084

Observation 3f825f0d-913a-4c73-8220-2720c572daeb · outbound

This paper cites Google scanned objects: A high-quality dataset of 3d scanned household items.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Google scanned objects: A high-quality dataset of 3d scanned household items

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.410446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8b71a293fe68db70fc9d104da81520499a5fe4fe27f64a8bdf6e0545ee460bbe

Observation 8f93d1e6-83f0-4f7e-943b-9d7a023e0fdd · outbound

This paper cites an unresolved cited work.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-10T22:58:52.413748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7b3a43a6ba84363a0ac904ce4d8f3851c1a23e1452ce33633da9c571f6d1c974

Observation d992ae0f-c056-43e8-8e4b-a13887ea7fe6 · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Taming Transformers for High-Resolution Image Synthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.084829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a2780ba0599d1f4c039dc56ab3c89bda4e376d546296e6f826534941ee4e861c

Observation 9f6ebd11-36e7-4089-9f9e-849e4be37272 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Structure and content-guided video synthesis with diffusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.417956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:44a48eba809e1b896b3f25e572a326eb634c9ac3f3b5d594628f5d4d98a0240f

Observation 47afb22e-77d8-4081-a419-162fe74e5bef · outbound

This paper cites Two-frame motion estimation based on polynomial expansion.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Two-frame motion estimation based on polynomial expansion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.423853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f1c2f090be53098287706e29c05df1a907b186028d7e8a243f9af3e490b80fe8

Observation 5c183e3c-648f-488b-8aff-7287d2f187b1 · outbound

This paper cites Stylevideogan: A temporal generative model using a pretrained stylegan.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stylevideogan: A temporal generative model using a pretrained stylegan

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.430349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9be8a2ebd7829dcffc998e8ae96a05bfeea52681fac9a4c24ac8df283c2e0618

Observation d13f1608-b9be-4757-a8c6-c7bb497ae565 · outbound

This paper cites Stochastic latent residual video prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic latent residual video prediction

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.433370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:0194b7c66e949c2f0a16cb19c17cb099687bb91f7003ed7ad5d064a1d4e983ff

Observation fcf8153a-93af-4fa2-927b-05d260211088 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.107162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3255949b45cac4f2f1abd3f36ae48f9307a678f44b2e43de58ae3e61e6a5c28e

Observation 690678e9-dadd-49a1-b826-a3233091cbef · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.440342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d562fb478eaf8410a4cf06eebe954eaa2e0ee2c1eef77a8a6227453f3b5f877a

Observation 9a39a536-4f21-4df2-ab0a-6b7b301a3307 · outbound

This paper cites Preserve your own cor- relation: A noise prior for video diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Preserve your own cor- relation: A noise prior for video diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.443748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:29a3d92537c553d1931a466829208102e7906b9234188ee15b0573fa33c80049

Observation 243d2d62-1813-4d4d-a44c-b8a5abda79b6 · outbound

This paper cites Generative adversarial nets.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generative adversarial nets

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.446463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:dcd522ac9c1c29e8aba5cfd4ccd0aa37ab6e12e2072977b4d4d7a7546bc9a92d

Observation a5b19aa2-1efc-46ca-bfcf-feb54b55f7e5 · outbound

This paper cites Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.031672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:176fc1e49cb4e8227689a534a3138c7dbd994ce66ee8707fea70f1d3d4835b6e

Observation e1db8544-3368-4414-b1c3-2e9bbd7852bb · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.081067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a5ed64f6220ea1140fa34bd08455c343dfeef3e5b36544bbda55cd52f948b9f8

Observation 2d998d1c-bee5-4b8c-af61-db9795cfdeb9 · outbound

This paper cites Rv-gan: Recurrent gan for unconditional video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Rv-gan: Recurrent gan for unconditional video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.449183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1f2493a73d9ef8d2f10c159ee5df8aab55b5fa916d53459af8de153994aeb7fd

Observation 91f40faa-f813-4d60-a401-8a1b820a41b4 · outbound

This paper cites Diffusion with offset noise.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Diffusion with offset noise

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.453606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6a2327a8133cb23f0073cf68b447325892553470088004755b0d101baee57541

Observation 7a87548d-6f96-4e28-a5cd-86b861b1dd53 · outbound

This paper cites Latent video diffusion models for high- fidelity long video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Latent video diffusion models for high- fidelity long video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.456584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:5fed77cee1fb975473346336e7d1990b21e8b99b6ccf5a28cfb2b08ed539d36e

Observation 47f51d8f-25a7-488a-ad22-bbb4152a5a59 · outbound

This paper cites Classifier-free diffusion guidance.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Classifier-free diffusion guidance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.459221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9ce3d777ecf186f9c2d5d6016fb368bee743fc3a96dc96b255dc22d644ef5820

Observation 2552e170-3aef-480c-b3d9-88cc076c19d3 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Classifier-Free Diffusion Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:51.924296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:991c5870b65efff9a73312464dfb7f264e4e13993692f2837edd401db9b3eadd

Observation 93c36309-9772-43af-a905-4895897a2b09 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Denoising dif- fusion probabilistic models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.461828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6fc25652d3fc1ef10b46c20483bd15da448904f3c88994da7445fe91082e8789

Observation 063774ac-f08c-4560-bad1-c08ba661c7d8 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Cascaded Diffusion Models for High Fidelity Image Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:51.972482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b7cfa87f7fa2d14ee4c2c2887e93177e07b73eca3ee03657849923158f43d181

Observation bac83294-60f9-4168-9ed8-008d896ab7b0 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Imagen Video: High Definition Video Generation with Diffusion Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:31:08.340948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8b098ec6020e6b3362e5b9b0769415919d1d8ac5cc45019d29cd17e62d0f1c96

Observation e481fdc1-9476-4793-8abc-787900d4fca1 · outbound

This paper cites Video Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Video Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:38:28.190773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c47fb44393d0cf00fecb8279254e4a18db25039118dfcf38e8de61c961ef3dbe

Observation bfa7556b-0994-4e58-9826-50e9a7be80cd · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.465451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:879e81eb36403f2e8121ad4edc7f6994e40d6ddd1ec83a643fa560ca13ea8271

Observation a2d7b100-b1a6-4787-b0c2-88e567d362a9 · outbound

This paper cites Simple diffusion: End-to-end diffusion for high resolution images.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Simple diffusion: End-to-end diffusion for high resolution images

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.051355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:0e5d8a3bfcda2e01f95504fe47e8b2c11bde209a5c3aa815abfc41a804e2352a

Observation 80ea7579-f424-4acd-99ee-ffcb0f3bfe4a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets LoRA: Low-Rank Adaptation of Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.070238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:60b096f19b5bffa4eb7e540dc1859e9c97710860624102925539d66d156658ef

Observation 3086b5be-cc65-4653-892e-e2273ce0e237 · outbound

This paper cites Estimation of Non- Normalized Statistical Models by Score Matching.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Estimation of Non- Normalized Statistical Models by Score Matching

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.467988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1a1154ec1a5b1999052b2ad0790068b7e82c62cb9c4651327d197c8629a35b9b

Observation ba979575-0310-45df-ba27-f1508672c505 · outbound

This paper cites Open- clip.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Open- clip

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.470262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:fb9a21464c8425bea5744bae4923399cbd6dfe082aa9bec55b52683fc56337ba

Observation f113f498-f494-41e8-8c7a-0e2fb4feba54 · outbound

This paper cites Open source computer vision library.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Open source computer vision library

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.472725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d90ef0fcfe829bb8ddd930a324a0a3b919817ec4bee05ae421c93a13bd1e6716

Observation 850225e8-c922-427e-b7ca-a5de58c892b4 · outbound

This paper cites Shap-e: Generating condi- tional 3d implicit functions.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Shap-e: Generating condi- tional 3d implicit functions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.475372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:27ed28a794887972f2d5b9e144a3fbd409d1e23a697953d0ff632672f6c9f350

Observation a244474e-b377-4cb7-8994-d8f16a92ef16 · outbound

This paper cites Lower dimensional kernels for video discriminators.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Lower dimensional kernels for video discriminators

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.478776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3dc821e7c9e255374553f84ac6b22e477c90dc35fd539dfdb5fab0f5a64e521e

Observation 7a1154aa-5e2d-4727-8b2f-4426575ecb7d · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Elucidating the Design Space of Diffusion-Based Generative Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:46:53.453910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:df19c72e6cddf2c552775b5756f26bb3d62de2ecb41636b24da1ee895800630b

Observation f98227f2-8af8-4cce-a16a-8e71eb6f1bc0 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.481343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8b96cbd79d96ad93e44d526ccdb62fc97ca6c268ed85b61c9d82e6ed63c03f9c

Observation 0b4bce9d-a8b5-4b0a-90ad-b1ee75f5da75 · outbound

This paper cites Variational diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Variational diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.483561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:47033b884441f6cdffef376bb8fcf9674ac19188f68185a59298e45158a5feeb

Observation afffa12c-1938-460a-9ce6-adc9320f50a0 · outbound

This paper cites Pika labs, https://www.pika.art/.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Pika labs, https://www.pika.art/

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.485890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f709fcb21f617aca34a7c7e9d81f579eb5b4c22413c8f026d9c52051b044c2db

Observation 3fca2898-f60b-487f-af8d-341ceb1f05fc · outbound

This paper cites Stochastic Adversarial Video Prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic Adversarial Video Prediction

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:51.976482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:5169e79d635412a196ae3cae6bd59acf0d39358aa3e0e417449cad6a674c4b91

Observation 3551b37e-644c-4587-ad56-bb649e380c8e · outbound

This paper cites Common Diffusion Noise Schedules and Sample Steps are Flawed.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Common Diffusion Noise Schedules and Sample Steps are Flawed

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:51.982261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:cd6f8a82668e44d8a67ca124d9e24a5ee3e6696cdb4494e21bb7ae740f30b13e

Observation 3dc4f274-ad4c-421d-98d0-83360ec06777 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Zero-1-to-3: Zero-shot one image to 3d object

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.488826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:284e4a251482e1be851fbe2cde0c46f5c38ccda42ca7a79564900dcaa0f91ec6

Observation 4f7f2d8c-d0c4-4c54-8d56-1dfd30365503 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:35:43.296571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:21654fb9c3c92ae8fe1c2c5906fa551dbb9764f6be4e675e09530107be8f0528

Observation 0345f43f-3b3e-4fe8-8b28-1b23a2831069 · outbound

This paper cites Decoupled Weight Decay Regularization.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Decoupled Weight Decay Regularization

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.010731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d554b68d4227027a4c8aed510d2e42b364604dba58e7768de9d54ca4e39519af

Observation 70223a73-0f29-4875-90ba-cb84849a1dbc · outbound

This paper cites Transformation-based adversarial video predic- tion on large-scale data.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Transformation-based adversarial video predic- tion on large-scale data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.159200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b7df7ff91b6ff1a309a7ee10ff2becaa45a65e7371ebb4f0aa3e35bbcb2e5a4c

Observation 10ba17ca-7d1e-4955-a69d-afc3eb79366b · outbound

This paper cites Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.162066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:85424d83d7a35a5bc4d851d85714690a7a1664b3efcd55e5e243f277d6f50b24

Observation 786d8902-6c1f-4998-8440-9e7b4ba5be80 · outbound

This paper cites Point-e: A system for generating 3d point clouds from complex prompts.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Point-e: A system for generating 3d point clouds from complex prompts

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.165074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d199d797601fb86631b52585c8656e7791de073746e16f1eee629ec96ee06829

Observation 11226907-7541-4eef-8724-e5170230b98f · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:45.999902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:e30dba04832116e163ad39f819275bedb98d316b68ad068a2bd6bebe422e138e

Observation f9985714-33e4-4fea-a5d7-d10d465cff3e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1b0d0796b4fd8a94ac109720d4e7178a5f5208c0db177be347cea863705aebfc

Observation 7370bdaa-45ce-411a-8e06-5168591e9001 · outbound

This paper cites Training contrastive captioners.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Training contrastive captioners

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.167803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:714af34a663703a637c554236b5899972937955d18224613cbe21f51ecf1c374

Observation 7d7aaa5e-a0d0-480e-9ad9-22ff8f14c19e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Learning Transferable Visual Models From Natural Language Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.074775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9b387fbc70e300eaafbca7b34d1c425899010fd1ccb242e7e3cf6829ec1e568e

Observation fb34dd4b-1bda-4d08-8db5-3bbd2325aea6 · outbound

This paper cites an unresolved cited work.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-10T22:58:52.170395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:44c662a0c1faadf77677cd798ac78f4ca53b865b761a334575b0979781a8eede

Observation a901f9ff-a6a0-41f0-bdbc-c1991e011740 · outbound

This paper cites How dall·e 2 works.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets How dall·e 2 works

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.175437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:db9688b2a8a3c72bf0855e91c7813672d4414039e9c308b2a1e51e83e55a132f

Observation e428c13a-b90e-41af-9ee7-d6c201de1a31 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.092030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a672c664685fa826cf8629989756a1df73e929b0cf295ffaa26a3e42d407f54d

Observation 2f8e63da-c358-482b-9917-e97dcb5efc17 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets High-Resolution Image Synthesis with Latent Diffusion Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.096430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:be83cb223bcb90f415f21aed34fe3f8d50d39819c521ee70a27d7ea86787c369

Observation 2a66a5ed-5b6f-4efa-9d32-435f2b17b5f2 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.103805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:fa75522eb93f47ae1442a8f7584cf28cc4c95fd586438372398a2b28f0ae91ab

Observation b7c509d8-9f86-4d8a-adcd-54bcf546cf10 · outbound

This paper cites Gen-2 by runway, https://research.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Gen-2 by runway, https://research

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.178198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f612b432e0ed42acce2a8c919ee570900e6c746dd30132edb70f7e131a14b688

Observation 5ed6e231-1a11-4c9f-9a22-312c2450f838 · outbound

This paper cites Image Super-Resolution via Iterative Refinement.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Image Super-Resolution via Iterative Refinement

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.111625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:34ad6e3eb26bc03d9a268ef4fad9ac8203f7d80ed6ef8b871a4e65fbdd79c7bf

Observation a642a08a-88d5-43fb-b08f-f323c26ee55f · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:38:54.138183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9694918cfc28f161afeaa55f7ead7d139447aec3e670263fca2b87cd98620fd9

Observation 8c4286f9-9084-4c24-8a4e-97f72bd68b88 · outbound

This paper cites Tempo- ral generative adversarial nets with singular value clipping.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Tempo- ral generative adversarial nets with singular value clipping

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.187263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:ab536b69d400f1dcd12229995d335b9f60eeff7fa645499018a7784db01fb25f

Observation 41416f5f-b999-466b-91a0-e07ec90dc987 · outbound

This paper cites Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.190766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:40a3238b40e8d2ac853889b843ff2b1bb8e53e485831fc2a1a655d7b3d06de38

Observation 4ab4e3b3-883d-4fb5-a4a7-c28d21fae26d · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Progressive Distillation for Fast Sampling of Diffusion Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:37:44.815756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9fb9841866b91945f45c33d77dc7552bd446c06b06d54e0095f7c422232d6525

Observation c0620f95-20b6-4d95-8a0d-5ebb612dd05b · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.195095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:2777d16d348c7500699fdc420382695f41885150a8e4fcbc698aa1798dc4f952

Observation 34feb37c-b32b-4396-a56c-7092ec300e26 · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets MVDream: Multi-view Diffusion for 3D Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:36:20.266102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:10a7d92ebaf0106fc40fa3073a41d1d2103add822604326e58ff23e80287f289

Observation 9889afaa-6ea3-4723-a270-acb42ef8f891 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:13:03.403852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:0e63afa3fc2f375657b6d359b48357628b0a155719aab86e967ef44ea3a2b943

Observation a2c8a657-23a3-4415-bf59-d6f14e87f27b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.198584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:13f0eb4626cf8f7ed4ed50f9a0734fb97085d421c207970e571e519b5cdc7722

Observation 49a95cb2-cbf4-435d-9106-4b640c78c46d · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermodynamics.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:12:28.687532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a232cd9afbabaef978b31650ae451792e1721a66a57afd784121f49db0eca5aa

Observation e14a40b7-92cb-4361-9d15-4735feaadbe3 · outbound

This paper cites Understanding and mitigating copying in diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Understanding and mitigating copying in diffusion models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.201576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a5e7a4d5ec5960c78bbd434fbf68ab545413fd04d502374db44231d47d9d4f6f

Observation 94f7f97f-5ae4-4be9-a194-770bd2b6d255 · outbound

This paper cites Improved Techniques for Training Score-Based Generative Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Improved Techniques for Training Score-Based Generative Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:51.954301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c4933dbe80e4e14833d34682011115c2699b71dcaf3bed8642bb835bcd8212dd

Observation 7e794e55-785f-4775-a471-dc44b8ae111e · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Score-Based Generative Modeling through Stochastic Differential Equations

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:51.958806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:e85e0daa69862806d4807458f6d86451e85ae921c0c28c9d281b3e99e0fb34ae

Observation 12eb4b0f-eb97-4a00-b0c3-e577c4f0923f · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:25:00.378223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6f9953ebf4223389ad1d93f34389d47f4f8b6454b9524ff2f2f7c1cdc251bb94

Observation 2188a425-4618-4d2e-a17a-c9f8587e3775 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Raft: Recurrent all-pairs field transforms for optical flow

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.209334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1cfbb99e200dcd3897580716c5330ce4c5c3c74c6db7b1f902acc1bca92ce7b7

Observation 10c673ff-2f67-46c1-b8dd-1f2ce189ba37 · outbound

This paper cites Metaxas, and Sergey Tulyakov.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Metaxas, and Sergey Tulyakov

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.212551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:abc8febfbf725a59319264a1f171d0687605f3a0b0ed60b25f8a441d70a7b3ee

Observation 1f66e2b7-e935-4535-9e1f-82f4c0c19672 · outbound

This paper cites Converting video formats with ffmpeg.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Converting video formats with ffmpeg

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.220100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c80544a92aba9a82b5188e8697d818e990d83b7b5e384692305eab8710b2a854

Observation f5a396a0-1a76-49d3-bc6f-7c805e91a29b · outbound

This paper cites Score-based generative modeling in latent space.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Score-based generative modeling in latent space

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.230113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8a3d938630d9b8cb83c1c924c98696f5df23436e5d66e87f63d0b2ed7b30de36

Observation 1bf1c45a-fbe8-4a90-bbe9-9f7fd079c56a · outbound

This paper cites Decomposing motion and content for natural video sequence prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Decomposing motion and content for natural video sequence prediction

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.234644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:79c5d7687963457f9e49efada5daf6642ef6ea996cd9d7f2a73be623c627403f

Observation cd012fbc-e7d0-4450-85b0-cd7b2ad9d52e · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:43:34.268786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:77fb7710157e0f0d22f91a5cace0ab16afd90bda5a343c099c158d53133e1f76

Observation 301c62a7-1e1c-4f9d-b3af-a1b3cce9a07f · outbound

This paper cites Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.239876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:2d0f72fecd1858d5c31dd74fa49b61013c069c067942b402a38b51f6eede4acd

Observation 0ede45f9-e289-4435-a4e4-d06e131f0308 · outbound

This paper cites Generating videos with scene dynamics.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generating videos with scene dynamics

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.248191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:21e98adecc8cb1537e824eebc4ce06f685e8d4ca3b2afa6726464148ae9ceef6

Observation 867eb3c0-1e04-4c90-b25a-f9635e1a6e03 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets ModelScope Text-to-Video Technical Report

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:260b19a6cd92782a05ed3d1646150037a1096a4ed28ed5020f79f8f8e200d918

Observation ddb335e0-703c-4c00-9b86-35e1f131e0e3 · outbound

This paper cites G3an: Disentangling appearance and mo- tion for video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets G3an: Disentangling appearance and mo- tion for video generation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.252462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:0ea5041d36159f4a05ac4d6050c3ad9fa605945aceb0751dca85550a9131ecfd

Observation 16e5198c-4583-429d-bb8b-1eb3433d4008 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.038771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1392117b573f2b801eb4d0ca820e5a5ea0122d3f963d9c7cd8b5dcee46645818

Observation 53774ddc-6ab3-4442-aa6b-2eb8bccb4cae · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.258107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:876083a288c706bb7bfe1d681488eb7ec568dd5a50f6801be1cd8881d2d24ddb

Observation 3deb0188-2f50-4613-a69e-4b1ac2472a23 · outbound

This paper cites Novel view synthesis with diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Novel view synthesis with diffusion models

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.262032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:778f48e9b0800378ddc962c1e946675f7db5bb8f5d16210ee9a689bf0a5f21b4

Observation 052d33ce-9e4a-4aff-8ad6-a4f078eaff1d · outbound

This paper cites Scaling autoregressive video models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Scaling autoregressive video models

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.270162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:e9023894441047b16e1f7978b7efaf8803f116839d30b2b4a4149223efce3ad3

Observation 498e28fe-59c7-4351-9027-343899b4de02 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.067091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:24b1b1aa98b8af814888cd27b0f0563d13f8d5ea8c4e1b1020242678de0ca3a8

Pith citing papers

Observation 677f517e-ed14-4944-b95a-9fa4b713fb03 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.543020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:34ee9894790a6f477b165e5d9ef524a54c0d7ee2b6de37ff43d61a9135c82d03

Observation 75e31660-5491-4556-88cb-ccdecfdb3d77 · inbound

Latte: Latent Diffusion Transformer for Video Generation cites this paper.

Latte: Latent Diffusion Transformer for Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:45:35.812408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:45:35.754742Z digest=sha256:367ee423e84ef71839478c00bda05ba1d4c52d59bfd1ffe7763be37a3198eb51

Observation 1f3eaacc-c372-456e-9c1d-a941e73eca30 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.419878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:9f4fc6edf15cc7872979de0b56a9778019938b1bb5cae22ad9375ba65926ee39

Observation 321e56ae-80eb-47fa-998c-2420a872f1e1 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:43:11.088767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:8a79b70423ecd18d3ce59460840719090674d07d82037e09019c7fee949a3bb0

Observation 076c8b60-ee63-4f63-8fe3-f9764c5aca6e · inbound

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis cites this paper.

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:27:53.615646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T08:27:53.446686Z digest=sha256:1eb57be51f2b34195ce9129dacdedc1080606541d34f2caf860c3d3cb3173e18

Observation 6425987e-38d7-48ff-841b-678099452bc6 · inbound

CameraCtrl: Enabling Camera Control for Text-to-Video Generation cites this paper.

CameraCtrl: Enabling Camera Control for Text-to-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:06:23.575572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T02:06:23.410241Z digest=sha256:d3e5fa69267c94565765391773d24777bb84fe0141227191b77c69ffec2b00c2

Observation c828e4c4-3656-491e-84ee-4f08741443cf · inbound

InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models cites this paper.

InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:15:33.457880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:15:33.430513Z digest=sha256:3a29a962ccf7d62a009bd33b23685ef89567b853b959dbc0682ef50405951ab6

Observation a3382bfb-fddc-4c12-a12d-8968b9ab22e9 · inbound

CAT3D: Create Anything in 3D with Multi-View Diffusion Models cites this paper.

CAT3D: Create Anything in 3D with Multi-View Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:29:51.094198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T21:29:51.062396Z digest=sha256:cdd026cc733ce7504af291b271fb1274ca28511dd74b9695c288b7d43798a61b

Observation 0bf5dc40-2682-4de3-8de1-cb57f57c26ad · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.690990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:fa71d918d6a68aefcd51b5fee71de5c25987f52f8a024f42d1b5689f1cfa71c4

Observation bb8bf142-82cf-43d4-874c-aa0f3ee16a3a · inbound

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation cites this paper.

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:34:53.169615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:34:53.141898Z digest=sha256:afb96202f2eb7780fa9180ef03946b43039560106d5f5e004cc3ffa7869b2e0f

Observation a3cdd265-58fa-4aa8-9359-fb467f8c3595 · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.491420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:2b83ffd63c956c568bcb6a6e118c3f28113c412b09569c6cef8e17016fdb3b77

Observation 94ea79f7-864c-4058-a759-95fb8bf2e203 · inbound

Diffusion Models Are Real-Time Game Engines cites this paper.

Diffusion Models Are Real-Time Game Engines Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:04:44.211533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T12:04:44.162343Z digest=sha256:18d7e52c6581441d692ee5ff84791d8307f74b13b99a8994087c05380616cf24

Observation 20e19241-9f9d-4c4d-8186-82bbbe15b687 · inbound

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis cites this paper.

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:59:03.767446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T22:59:03.642189Z digest=sha256:7efce5a90340357718140b9e14e2ed51f1caf5113449e50dcf2843ee0f2bc76f

Observation 83453ea2-74ea-49e4-aa2b-fcc77435730a · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:56:06.835877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:423ef76ba6f930444927fce215822739f4672c3883882cd825363d831363f0f3

Observation c0cdda6d-db31-45f2-a834-168fdac0b0dc · inbound

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think cites this paper.

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:09:37.066514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T15:09:36.982610Z digest=sha256:7753c358d25bc1c67b3220f9682572a88cb6d15e666fdd1375ae083e0d1192b3

Observation 10a9d913-aaf0-4b72-8189-aac96663ddfa · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:20.818910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:98f33f0b16cccbf484b844090b1a3b99a82aa5de5811fe502b778d0d39714472

Observation 84e4c6b2-ee04-44c5-9037-1d6f79041a67 · inbound

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation cites this paper.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.389473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.389473Z digest=sha256:42deb96543d26cf749e2476ff3508df72ddd6aed4128bde767cdb9d2a54b062c

Observation 565a1922-dc46-4826-8cb3-3c7397282881 · inbound

Inconsistencies In Consistency Models: Better ODE Solving Does Not Imply Better Samples cites this paper.

Inconsistencies In Consistency Models: Better ODE Solving Does Not Imply Better Samples Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:18:03.727266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:18:03.727266Z digest=sha256:9b34f996384b9ddad0677c9350d6ade1f081fcb11fe012ff2f37a50474cef6fb

Observation ada0954a-fa50-4a24-8d99-61e35478ad2a · inbound

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation cites this paper.

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:04:21.804079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:04:21.804079Z digest=sha256:a4237570e37cfa454f409f876951c550cb38b10dfc470050cde677407fcbe30a

Observation 5dc5de44-5a52-4d42-92f1-7ab9ecf08cf9 · inbound

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey cites this paper.

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T20:53:14.800657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:53:14.800657Z digest=sha256:ff0d8f0d68d979b5cd9572925a6c4bcacde32d8b11e287f53ae8e3f644aad6dd

Observation 11152d15-d8fb-47ee-80f7-4bc5d27a4c12 · inbound

FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations cites this paper.

FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:20:48.099378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:20:48.099378Z digest=sha256:4e9b551dc841dedc480ca79319d6c6d402405d351703b222a113774fda04a192

Observation 393f26d0-2319-40c7-a8a0-6d1ea5e00796 · inbound

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer cites this paper.

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:56:47.843475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:56:47.843475Z digest=sha256:01c6d349f9b82146f8f19a0810732a87b5593744750007e4d2538da57ef2cd88

Observation 1603c780-1067-450f-bb66-76d1021955db · inbound

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation cites this paper.

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:32.332554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:32.332554Z digest=sha256:6e08d3f0ba745193a6fb582da35fca6c26c4a20a84607b22adf0b3610859ffdb

Observation cf62f43f-c85c-4500-b4b8-88f420ad6239 · inbound

Generative World Explorer cites this paper.

Generative World Explorer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:09:24.021530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:09:24.021530Z digest=sha256:0fcf29ad5a0767abe2a779596bf835794bd607370d23e9805b86b58306fe504a

Observation 52e97db3-16c8-4758-afbb-0b35f933e8a7 · inbound

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input cites this paper.

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:23:28.365169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:23:28.365169Z digest=sha256:4f8f5f24b9b5a3e08db45e77cd941cdbe20201d04aa16e3d501532cf5c844fb8

Observation ead9e6a5-5e03-493d-b160-f594c6da1e47 · inbound

Medical Video Generation for Disease Progression Simulation cites this paper.

Medical Video Generation for Disease Progression Simulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:11:36.525393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:11:36.525393Z digest=sha256:a5159157a59ffc8c994342bb2d834ce18e43906cb64a091bfefd713eb9c01072

Observation bf4effda-49b5-4da8-a1c9-b578898fd42a · inbound

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting cites this paper.

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:34:25.845410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:34:25.845410Z digest=sha256:2266292911042e0c0a6717eb807126e9d94709f5af973e2f80f1c883631ac966

Observation 160307a6-ca9f-4acb-8007-32a238f273ea · inbound

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models cites this paper.

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:43.958215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:43.958215Z digest=sha256:16adde65f2727268d9e1f2421f4a769b79788d7fbbe243d55e9efd4427fffabd

Observation 7ed2eb50-bcf6-42f5-9b9a-992f78905002 · inbound

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos cites this paper.

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:48:12.512254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T16:47:09.717923Z digest=sha256:c108e8887eb916869d5d2be955f860098f94dca6b57a2443f8b9ee029437a73d

Observation 29bcaf08-54ce-42b0-b31d-dcb75b4451cc · inbound

What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality cites this paper.

What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:06:37.364487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:06:37.364487Z digest=sha256:702a4f6e870c3461871258842d207b9394c1dbcf3f4e9c0811a144ca9d4c6076

Observation 38660f6e-eaa7-4cd0-a368-71b6c93ced1d · inbound

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control cites this paper.

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:56:27.063283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:56:27.063283Z digest=sha256:c2bc989686814a502393f222a6fed4e379647ff14573a52e3e9528eeb28ede87

Observation b3017906-ab9c-4d57-87a3-9f90f9b6e2e1 · inbound

Transforming Static Images Using Generative Models for Video Salient Object Detection cites this paper.

Transforming Static Images Using Generative Models for Video Salient Object Detection Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:44:08.089281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:44:08.089281Z digest=sha256:c46d89003fa2f8e54f8de0152371d8da75334ddf1b8b6278e283c2474b954281

Observation 7ab1391d-a32e-45b7-8de5-8c2d27109c60 · inbound

Novel View Extrapolation with Video Diffusion Priors cites this paper.

Novel View Extrapolation with Video Diffusion Priors Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:30:23.336641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:30:23.336641Z digest=sha256:99e3ed9f483bae3fb49624e3264c14062931f943a2b112aad255d61f0ff69913

Observation 6d963a4e-9d0c-4a52-89c0-76b0c9c2e4c2 · inbound

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation cites this paper.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.528135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.528135Z digest=sha256:0e15fcb3f6dd7bfbd375018e3780ad7099a0714891f5f02ce930f9b7e4ae0a12

Observation 63de849b-7c06-4ff4-8926-169e3a8dc27f · inbound

High-Resolution Image Synthesis via Next-Token Prediction cites this paper.

High-Resolution Image Synthesis via Next-Token Prediction Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.648696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.648696Z digest=sha256:c6e0f3ac3db582440f95d46197357a05063a3a5cede36784bd3d6453cc007712

Observation a68c16b4-5018-4766-aa85-77a91a44ceab · inbound

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement cites this paper.

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:25:29.354394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T08:25:01.468957Z digest=sha256:bb412d22b885215623372cd7287567b80af7d5bc088d50fd291f4e6e960adfac

Observation 9116f29c-f09a-47b2-95b3-3d823ad99807 · inbound

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation cites this paper.

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:11.542684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:53:11.542684Z digest=sha256:0955178c161f52318fc4456c9a614b8d21fb46bd7cdb80fe0d013e99b8281dc6

Observation 8374e394-8fd1-46ed-9b7c-025d35695774 · inbound

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model cites this paper.

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:34:56.915543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:34:56.915543Z digest=sha256:5abafe88a8cdfee78ee99e0e0e0bc343d94148ec11de919d38ff12c820df854e

Observation c9223a25-2189-4a83-8888-b3051444de5c · inbound

Sonic: Shifting Focus to Global Audio Perception in Portrait Animation cites this paper.

Sonic: Shifting Focus to Global Audio Perception in Portrait Animation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:17:46.427178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:17:46.427178Z digest=sha256:75ce20dc89456753b2123923c225dc531761b154019bcf6b492cf086cf47c530

Observation 9f6ae92b-648b-417e-bf98-e2782f8b95d1 · inbound

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing cites this paper.

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T13:17:41.855795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:17:41.855795Z digest=sha256:199d5a921fe6fb8e818e23388d777a6e851b8f37ef81e2887d88eb7222b231f3

Observation 0491205b-7fb1-4233-9078-09704480f6d0 · inbound

Privacy Protection in Personalized Diffusion Models via Targeted Cross-Attention Adversarial Attack cites this paper.

Privacy Protection in Personalized Diffusion Models via Targeted Cross-Attention Adversarial Attack Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:10:06.136789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:10:06.136789Z digest=sha256:f438f7fb5c2e3e35d83b8b1b60aac73ab79420a0f8be535d918d42a2c12b97e1

Observation 7346384f-8588-425c-b6d3-2dd8d8fc40d7 · inbound

Generative Omnimatte: Learning to Decompose Video into Layers cites this paper.

Generative Omnimatte: Learning to Decompose Video into Layers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:56:32.798127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:56:32.798127Z digest=sha256:8d85fdcc969f511a82423702075a323202cc50b08728f52c9b965fa96196e591

Observation fdc87925-0efe-4256-8b08-76cf1c4d6812 · inbound

Importance-Based Token Merging for Efficient Image and Video Generation cites this paper.

Importance-Based Token Merging for Efficient Image and Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:43.242671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:43.242671Z digest=sha256:0bdb67ea9fcb78cecfa2c6e197d797db6ffa0ded11370cfe8961efa334c8a8ff

Observation 0c381b0a-75ac-4c4d-90d0-debd072b25f5 · inbound

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion cites this paper.

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:20:34.546964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:20:34.546964Z digest=sha256:8360272e0d7d1f6998cabc48ca18181c1418032ccf7f676b36b5b08bb6204c94

Observation 0773e211-3941-434e-90cb-5149ee17a483 · inbound

Pathways on the Image Manifold: Image Editing via Video Generation cites this paper.

Pathways on the Image Manifold: Image Editing via Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:03:42.363176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:03:42.363176Z digest=sha256:08bf19276048610418357120c086f6acc273f796d39029a5269430b5a3b3fb15

Observation e99f3425-8652-4e6f-b561-f8c14e98878e · inbound

AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM cites this paper.

AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:27:41.518760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:27:41.518760Z digest=sha256:fc7d5ffec99546f217cf31b1f3469c4274efb12e1e58f99ee40f5c42b3f98fe9

Observation 2518f65c-8653-4fde-acf1-c5c1fb3e8f37 · inbound

Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors cites this paper.

Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:26:49.345431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:26:49.345431Z digest=sha256:6080547026794cf69a555d3354a05fe0e93a5a9c228cf6b1741ea3424ebe3f37

Observation 848a662a-caff-4f38-9a73-a640104cf219 · inbound

AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation cites this paper.

AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:18:02.480654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:18:02.480654Z digest=sha256:1cacb2d0f32336db3573699cfbd10024de34ea60c64c5e087b9cdbbc3bbb0b42

Observation 5a8391a1-cc4b-442d-8d99-b9af9bfdf7a0 · inbound

Identity-Preserving Text-to-Video Generation by Frequency Decomposition cites this paper.

Identity-Preserving Text-to-Video Generation by Frequency Decomposition Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:27.167230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:27.167230Z digest=sha256:6d688a58fbee96b0f43fadad12cf3e620b4bb6194275dc48d7b7f460a3994f82

Observation ca0b4a33-3d29-4932-be9b-30eabaa338bb · inbound

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model cites this paper.

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:11:26.703089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:11:26.703089Z digest=sha256:f19575b3185e0b667b9bb87de47c620998ad09fbf5725cda40f15dbe48497099

Observation 523ba06a-a3aa-4b23-ab8f-f07675bb14a4 · inbound

Towards Precise Scaling Laws for Video Diffusion Transformers cites this paper.

Towards Precise Scaling Laws for Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:57.085344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:58:57.085344Z digest=sha256:8fcb1a59868c12d576e390fb30870d90fafac687fe04d4188a9a70f47da90c2d

Observation 8314ebe8-b45e-4a36-8c70-8b77bac7a363 · inbound

StableAnimator: High-Quality Identity-Preserving Human Image Animation cites this paper.

StableAnimator: High-Quality Identity-Preserving Human Image Animation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:28.113601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:56:28.113601Z digest=sha256:31774045e210866361b74c5abeee8a9900e3ca1e94792e9ce6df3199ad4184a9

Observation ae14bb78-e9c3-45e5-945d-5ea7b1fbbe2f · inbound

I2VControl: Disentangled and Unified Video Motion Synthesis Control cites this paper.

I2VControl: Disentangled and Unified Video Motion Synthesis Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:49.118636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:49.118636Z digest=sha256:a12a47fcb78663b788e077b69b380ce4ad7b996b7a653a5aaf0eb24193cbcf37

Observation 7cf75921-3476-4144-b84f-643cd6039b9d · inbound

HiFiVFS: High Fidelity Video Face Swapping cites this paper.

HiFiVFS: High Fidelity Video Face Swapping Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:24:09.449935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:24:09.449935Z digest=sha256:c33dea13a2d0ec65fe34ffb78e52dcb3a9f83ef8d155fc2f092df938113f70c0

Observation b68b3016-ccce-483c-ba5d-265a41f3fa12 · inbound

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models cites this paper.

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:20:37.723736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:20:37.723736Z digest=sha256:9cdde99be69339a79f3a1b96656b5d6f3303ed028082a683b7b3f92618f05a93

Observation 94910b81-f451-43b8-8a45-a62c6d7f4d1e · inbound

CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models cites this paper.

CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:05:32.553141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:05:32.553141Z digest=sha256:918b19efd7dd6a09254fcc65132e965b80c2785dcfb63e1834685dbaee3f0fbe

Observation 429afd48-6d0a-4214-af27-db9bc1c5e853 · inbound

Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop cites this paper.

Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:51:36.457544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:51:36.457544Z digest=sha256:721540458d1bbeef60c21eb5c5d40e7815c93cd3c7b40694040f4c53e622e769

Observation dfc86daa-e6cc-4ebf-a9a2-a3bb04d3a2fb · inbound

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling cites this paper.

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:14:14.980281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:14:14.980281Z digest=sha256:f7a50c44148b78da5001a96f11e54f51362c214cb316c1126d299e5b58743f4b

Observation 6deccf3c-9644-4f84-8d6c-7246b6df6751 · inbound

Towards Chunk-Wise Generation for Long Videos cites this paper.

Towards Chunk-Wise Generation for Long Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:11:52.035699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:11:52.035699Z digest=sha256:5ee9715db089deaa77d08c5efe7a57323558abe42afdb3b691a1abecf987e09e

Observation 0fd80ce2-7101-4b02-863e-a416b4534ddf · inbound

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers cites this paper.

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:07:39.401420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:07:39.401420Z digest=sha256:1c72e52dfaef21b7303497d58b26749255d58e65aa15d31ae73ddccd1be03589

Observation 53844a28-1ea9-40e3-81ab-6fa1467c4ec7 · inbound

MatchDiffusion: Training-free Generation of Match-cuts cites this paper.

MatchDiffusion: Training-free Generation of Match-cuts Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:05:09.632876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:05:09.632876Z digest=sha256:92e58773634af2b95902dfb9f7c3f0b4f477389f97518e9a25cf9466c54def50

Observation f6c191ab-df5c-46c7-8c80-c8ccb6961f80 · inbound

RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning cites this paper.

RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:51:56.409783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:51:56.409783Z digest=sha256:82a90cb8783606006efac9f320da311413831b1d268c221276cc1e299ec8d64d

Observation eeee9981-fa00-474e-baa8-eb65bfa3d1f9 · inbound

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing cites this paper.

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:14.674984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:14.674984Z digest=sha256:067847332ea60abb0689f6472a59ffde1744989ecef5a3f2668da60ac1c6146a

Observation efca951a-66a1-401e-a52e-7e9aaa51eb96 · inbound

PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors cites this paper.

PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:40:43.371378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:40:43.371378Z digest=sha256:03b54cf48dd0a4c75399033deb6c1b586deb643d698f0ee555070bfa7ffd2d93

Observation 472b89b3-4891-4dc7-a2f1-4eb3cfbd4b75 · inbound

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model cites this paper.

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:25.003650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:25.003650Z digest=sha256:664c188788d64934e7ae78517e44bc09974874c2da2f2fccd8de5559a4c27411

Observation 75bbbc40-938c-48f2-979b-ed448ffca992 · inbound

Video Depth without Video Models cites this paper.

Video Depth without Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:31:01.634186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:31:01.634186Z digest=sha256:f39674a32d6f90e9e9188b26071ce95f7a90bc7e2b791309fe7547479fa5948e

Observation 85493b80-5517-42c4-b137-a042533dd9d8 · inbound

Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes cites this paper.

Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:25:21.832731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:25:21.832731Z digest=sha256:ae501c9edc09b95a7d607dbe3c44598eca4aa1be45e21da72b7c24339bf0b1ba

Observation a728a35c-4f27-4461-9da4-62b7f0fe18fd · inbound

Trajectory Attention for Fine-grained Video Motion Control cites this paper.

Trajectory Attention for Fine-grained Video Motion Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:21:23.856183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:21:23.856183Z digest=sha256:17a8945ad689dd7cf18dbe7593cabb6ced836835dc89ed71c3ffe97791374430

Observation aa37aba5-55c7-4dd9-a31a-5c6e062c5d5e · inbound

Fleximo: Towards Flexible Text-to-Human Motion Video Generation cites this paper.

Fleximo: Towards Flexible Text-to-Human Motion Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:40.490628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:13:40.490628Z digest=sha256:34ee853ccd71595d06d6f948e2e64ffbb624a57838c39a2134a01265a433c583

Observation 241544b3-e6b4-4dba-84b2-a965d4e59fcc · inbound

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook cites this paper.

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-12T10:08:47.608457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:08:47.608457Z digest=sha256:a75dd81e83c86e0c33e314c730a757e1f45d8c73fe12e930e75293a8828565c8

Observation 31993a5c-f1f5-46e4-93a8-28e6fdeacb1f · inbound

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration cites this paper.

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:09:30.510135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:09:30.510135Z digest=sha256:0cbabd1cfbb2784f8f348b4044887d8fddc298c53b7cef3a256b54e3b5b9707b

Observation c5458a79-4f71-42ff-b7b8-286731055181 · inbound

TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting cites this paper.

TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T06:03:03.692589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:03:03.692589Z digest=sha256:4fe2d9b4ae1f76ddaa9dc2d6d65d7e11636f77f6e567acd590408cf84087a19c

Observation da528990-3540-4961-b533-26120a10fcc0 · inbound

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation cites this paper.

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:52.396362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:52.396362Z digest=sha256:8d0dd5e49838149f09a4b22d497c180f421902fbaa646b3e83dce9c35dce6aab

Observation a49ec532-a065-4d3f-af8b-b838a26ca687 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.258752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:13bde54d780e02625e260ed9bb9ce6af7209dd84591c28be9541be32216c37d2

Observation 73abef5b-6755-4ee4-859e-20d4f321bf19 · inbound

Motion Modes: What Could Happen Next? cites this paper.

Motion Modes: What Could Happen Next? Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:15:31.831422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:15:31.831422Z digest=sha256:32b53a81d63d7805f3dc8c9f1fda3fb98808512aca3fc1cfff4b64862197a728

Observation 5e757bfd-07f5-46ee-961a-5eea5c882c7e · inbound

AerialGo: Walking-through City View Generation from Aerial Perspectives cites this paper.

AerialGo: Walking-through City View Generation from Aerial Perspectives Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:43.546739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:43.546739Z digest=sha256:86daa0d620599dcaeb58af301e0cc8766ee51c9ad8ae54d7949ba6a14c2d4d9d

Observation 72dd2002-bf37-4940-a96b-f8428733cf5e · inbound

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses cites this paper.

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:29:09.400491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:29:09.400491Z digest=sha256:026080c3bc257a924f61083dfc76c328f28bbaaefcee674f3ea0a4aff6cedb40

Observation 9f07d460-e61c-4e4d-bdc7-9dcf7f770f24 · inbound

Human Action CLIPs: Detecting AI-generated Human Motion cites this paper.

Human Action CLIPs: Detecting AI-generated Human Motion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.620799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.620799Z digest=sha256:a093c1bd4cdc4119706464077dfa572229e32802e243ffa83c8eb14f571bfab3

Observation 18988e7f-9dc5-404f-9084-cce3e638ef9c · inbound

Playable Game Generation cites this paper.

Playable Game Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T04:58:26.922490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:58:26.922490Z digest=sha256:4356359b9717d89929f0467bfc24cc89b4e85376083dae86d06027ef1cf788eb

Observation 6e2bab96-5313-4cb4-83c8-32003370b88c · inbound

One Shot, One Talk: Whole-body Talking Avatar from a Single Image cites this paper.

One Shot, One Talk: Whole-body Talking Avatar from a Single Image Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.204504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.204504Z digest=sha256:e722a1eedeb04b7ab80a7fea67e811de369ee046d683eeb02b24dd49081efd78

Observation 3db66a20-5065-4c40-85ea-92dddfad7f83 · inbound

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation cites this paper.

CPA: Camera-pose-awareness Diffusion Transformer for Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:23.825112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:27:23.825112Z digest=sha256:e1c2b3fddeb90fb7f72e14399623d7056f22a8ab2569bee0c8b939af5ed3e1be

Observation ec2399cf-bf68-446c-bea6-32c26bcd16e6 · inbound

World-consistent Video Diffusion with Explicit 3D Modeling cites this paper.

World-consistent Video Diffusion with Explicit 3D Modeling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.321343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.321343Z digest=sha256:3d5f9db6db50a11938da8d16c0b1c06d7c7f51032f05d9fc1bb709cb374082ec

Observation 746ffd17-70f9-4af0-a183-6a1a1a6b9442 · inbound

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions cites this paper.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.675581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.675581Z digest=sha256:2a343c50c38b903b2c5b7c722b252bfe306ab81fa1255995e4569c1fff241ea6

Observation 3a8530e8-25b9-45cc-b074-db5796b2cc5b · inbound

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis cites this paper.

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:28.574952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:28.574952Z digest=sha256:7f799735f217fc8879e28070e01d2d7dde1a2e9d3bbc9507c00c8d80adf8f60f

Observation c9f6f8e6-0976-42f8-8cb1-b705222ce6ad · inbound

Realistic Surgical Simulation from Monocular Videos cites this paper.

Realistic Surgical Simulation from Monocular Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:41:26.891437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:41:26.891437Z digest=sha256:027a5e31f09cdbde2f0b20b4218065179759d431ea66e9b1f70de3d083c060b5

Observation b3b54629-ec09-4ccd-b41d-4437a51e7b80 · inbound

Motion Prompting: Controlling Video Generation with Motion Trajectories cites this paper.

Motion Prompting: Controlling Video Generation with Motion Trajectories Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:09.508424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:09.508424Z digest=sha256:2023e3b57f5912a9e4f67b0ccfd03db324615ea3f4000b5b16f7830fab2ed280

Observation dc2703c4-4df9-446b-ae3f-c6d6f1b7fdb4 · inbound

Mimir: Improving Video Diffusion Models for Precise Text Understanding cites this paper.

Mimir: Improving Video Diffusion Models for Precise Text Understanding Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:11.634791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:11.634791Z digest=sha256:1136ec43df63b6358fb9ce54d3e618a615d59a64b91e842854f09cbd9dc3d766

Observation 27e71c43-362d-4814-bd33-65e155686682 · inbound

NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images cites this paper.

NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:22:03.731286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:22:03.731286Z digest=sha256:5fe3df851dd119780405619eee858c2dec288bb7ed5bc41ca1f6c95e67a75402

Observation 95dae0a5-a140-429e-bae2-3f7a17ec6b12 · inbound

Imagine360: Immersive 360 Video Generation from Perspective Anchor cites this paper.

Imagine360: Immersive 360 Video Generation from Perspective Anchor Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:23:15.397135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:23:15.397135Z digest=sha256:a29e21bfa12e5a27a55e3d7ddf9cbac06ceba2824c5288b3f3d5bd1d61dc0298

Observation 6e2bcb3f-4e3d-4f15-9319-b75e40814483 · inbound

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation cites this paper.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.477120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.477120Z digest=sha256:74a2341e9e2a8f4f56e376dba442376c3cd3317b2d11f5fa3d8e0ef154bddb46

Observation ff12859c-97dc-4101-a8e3-0cfabb908a13 · inbound

FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes cites this paper.

FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:31.010412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:31.010412Z digest=sha256:2beb1e56089f8837740f7c977577a3c5df7e1566211768779c3bd24b1055aa57

Observation fb0bca9b-a2ce-438d-81e7-e323d6e6721d · inbound

Navigation World Models cites this paper.

Navigation World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:09.705748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:09.705748Z digest=sha256:5b935f9e97b2abe81b21c3e9e7ecae5769ef47d5616e0dbbb74c18235712ffd0

Observation 19aa098a-4fb0-4a97-a98b-6034561afd07 · inbound

MV-Adapter: Multi-view Consistent Image Generation Made Easy cites this paper.

MV-Adapter: Multi-view Consistent Image Generation Made Easy Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:20:31.418758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:20:31.418758Z digest=sha256:1e123287e7f672f7a25269d80b6847b46175130dcc0b62f4f3be4015f54ca0e9

Observation ac079213-b637-407f-941f-fa609d26185e · inbound

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models cites this paper.

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:58:38.068259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:58:38.068259Z digest=sha256:347751f7a9176cbb0a7ed75ff4bc3e18bb24cca7e98845ff9a7a98955767d13b

Observation 1141abc7-39f4-4ee4-8990-e6794b58c03d · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.188344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.188344Z digest=sha256:7f1ee149ad76b4831fb36d9f8a5daba719bb62a91e9f69bdb8cea4563970e112

Observation 00949303-f71d-41d2-ab55-2af4b7c42bd2 · inbound

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration cites this paper.

GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:20.845667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:20.845667Z digest=sha256:e8658e4e33854eb5da24c735c2fb5ec186afb8bb2e8a991f3469694da14d1946

Observation 8ad0522d-d14d-4c64-bc72-0403f2d0a388 · inbound

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation cites this paper.

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:25.477711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:25.477711Z digest=sha256:03104adb1afe7694710594e02c76e39464aadf96ea7b6a908566955a0781e2a1

Observation c7633e65-5254-4364-85a8-a7e2253132b1 · inbound

Factorized Video Autoencoders for Efficient Generative Modelling cites this paper.

Factorized Video Autoencoders for Efficient Generative Modelling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.607910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.607910Z digest=sha256:6cc180c1848e3ca765a62c77b6d3bac9b85734b4663b7aee97ad3b63fdaf68ac

Observation fbe116fe-4c3f-4a21-82bf-136f5980f7a8 · inbound

Mind the Time: Temporally-Controlled Multi-Event Video Generation cites this paper.

Mind the Time: Temporally-Controlled Multi-Event Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:55:01.863979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:55:01.863979Z digest=sha256:0e98060c6a57f8f4192aeebc0d49276c58fd7791509112d8963c43af238bfece

Observation a334d119-9f6d-4022-9229-0a8b19a8ca4f · inbound

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation cites this paper.

MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:20:36.436348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:20:36.436348Z digest=sha256:f520c999d5c6dce169e58cfb46f3a1e7eb8e1fd57b5a5fe677fc4a7344a564b7