Pith. sign in

Paper Citation Record · LEDGER

Factorized Video Autoencoders for Efficient Generative Modelling

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.04452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04452 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:03.942929Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa86f80b-4a93-4938-b98c-d8b03c60c694 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.595252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.595252Z digest=sha256:a974fb448b44b80538258e4393e45350afea4e7f3ec33a2c64f5781400c977f2

Observation 6381c8a7-b767-4739-9705-e4a0a24b720e · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.602198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.602198Z digest=sha256:d01f538acacd4c3bf08051a5038aedf85543cc48aeca2de2608146deea249a5e

Observation c7633e65-5254-4364-85a8-a7e2253132b1 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Factorized Video Autoencoders for Efficient Generative Modelling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.607910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.607910Z digest=sha256:d06035b34622b4cf325deabfbc23d2698532f9c5e5b22b91d957d6fb157d1cd1

Observation a7104834-8ac7-419b-b46b-34f49acab452 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.613538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.613538Z digest=sha256:7f285222f547931d53340e8232464e573097bd9425db0622dcf9bc13a33a5629

Observation 252db8f1-7a76-4f7d-9459-7364a68e970f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:05.013267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.618977Z digest=sha256:0d11e4762d669a6eda4a1b4fbea85510585e2b0d106f7b6eb99b77b5a790c354

Observation 292a12cb-914f-4042-8834-1a702a8f5ff3 · outbound

This paper cites A short note about kinetics-.

Factorized Video Autoencoders for Efficient Generative Modelling A short note about kinetics-

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.625396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.625396Z digest=sha256:bf576ac3ea3e110cd5fbe90f0379a34c7789128433e48fda9a5cedbe1666937a

Observation 0da9b74b-b9b5-4420-bfa9-41350f215fa6 · outbound

This paper cites Tensorf: Tensorial radiance fields.

Factorized Video Autoencoders for Efficient Generative Modelling Tensorf: Tensorial radiance fields

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.636348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.636348Z digest=sha256:cd37e0cd503578102e22ff153f66fa5e92871e4bfeb15897c8c58918d3b63672

Observation 4d548e13-3cf5-4307-aa7d-a1ed5689ff5b · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.640968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.640968Z digest=sha256:70fe1a4297868dc7b0bd4c25cf5c8052908b3a5fb90983162ecde7d7a03c1e4d

Observation 919f90c8-9b29-4470-bcb4-e15df0c00b7a · outbound

This paper cites 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion.

Factorized Video Autoencoders for Efficient Generative Modelling 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.646167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.646167Z digest=sha256:676de525885931f1d18a3c0963d346b95bf83193ce877d611fa71e1f8428de8b

Observation 056dffed-8d84-4955-a67b-df1451abfcd8 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Factorized Video Autoencoders for Efficient Generative Modelling Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.651548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.651548Z digest=sha256:45402bbdbd1b2e25bcd060fcfdfcd4bced7b07b893b0388f536dfd7780602e3a

Observation 8ef58888-fdd5-49b2-81e1-ad6428c58a03 · outbound

This paper cites Ldmvfi: Video frame interpolation with latent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Ldmvfi: Video frame interpolation with latent diffusion models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.962108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.657483Z digest=sha256:c5f88682c7b65e6453136c564122fea72c3a042c96701a9922ad56db854979cc

Observation 8946269b-0a5f-4905-97bb-429daa074e00 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Diffusion models beat gans on image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.662268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.662268Z digest=sha256:dd43a5feb012df53c65f233cac5a26112f87ae9fc953054c1dd24aef789e4658

Observation 915a0a3d-2d8e-47db-ba08-c64866abea16 · outbound

This paper cites Video frame interpolation: A comprehensive survey.

Factorized Video Autoencoders for Efficient Generative Modelling Video frame interpolation: A comprehensive survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.932701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.667877Z digest=sha256:bdde329129c7b602a6a92b3df1c675b42cbfd185754f1454cc7fae27a940731a

Observation 4c1312d8-7e09-4380-b4ed-1010ae77c4a9 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Taming transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.672537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.672537Z digest=sha256:d24f247007158a27120b61b32144320d2c10905215c9a8ac2c3f96f3fff06363

Observation d6d89e8b-9947-4e00-ae07-5eaec08b4dce · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Factorized Video Autoencoders for Efficient Generative Modelling Cosmos World Foundation Model Platform for Physical AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.677424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.677424Z digest=sha256:36543a9811d9e846cfa0bcc9dff6fefe78d61f9a3a8e83cc0f8b8693c431f193

Observation 680b2c94-eb65-4f5b-aa50-da2e203ccc21 · outbound

This paper cites K-planes: Explicit radiance fields in space, time, and appearance.

Factorized Video Autoencoders for Efficient Generative Modelling K-planes: Explicit radiance fields in space, time, and appearance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.682506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.682506Z digest=sha256:ee50faab3e8de9fa7185601eef89d53ccac6dc3d984a05b5ac2aeb2782fb0c08

Observation ec7def8e-cb2a-45f5-80aa-6db2cea31c5e · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.687129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.687129Z digest=sha256:63fabdd0417f4c52d3e2b676defaf0c0c090e3d776de2c1e34022008b5d4f2df

Observation ada504c2-6454-43ac-9c76-7cd4c02c325c · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Factorized Video Autoencoders for Efficient Generative Modelling AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.692294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.692294Z digest=sha256:e5275fbd940c8859989a796f37a51c4ebe263d34673b174852d8f6286e94421a

Observation ddd0a472-6b8d-4915-97ba-73361a3e8439 · outbound

This paper cites Photorealistic video generation with diffusion models, 2023.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic video generation with diffusion models, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.892252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.697190Z digest=sha256:3dab8feb77c89ef913e27c8c71c8492707da05463d0719683b7158561e414356

Observation 3e6476b8-a059-42c3-99b7-b13be2bf68a2 · outbound

This paper cites Latent video diffusion models for high-fidelity long video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent video diffusion models for high-fidelity long video generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.702493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.702493Z digest=sha256:c2f689437dbefd2bf4907f828e87d0ee16feddbd62d67d0bca8b0ceaff7db1c6

Observation 3c6490a1-daed-47ad-8345-694d17d96753 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising dif- fusion probabilistic models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.863702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.707231Z digest=sha256:a0216275ebd556cb7640ee353eec68073a7cf512c3c1b4958a7e53c4e075907d

Observation 801cfd4f-9a6a-458c-a650-270d12715f80 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Imagen Video: High Definition Video Generation with Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.712385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.712385Z digest=sha256:77124a95dead4b9050e2502b66ead1c679caa95b174e60f4290f7ebc7863c957

Observation 78898a73-5e60-4ba1-9608-794b6c61d323 · outbound

This paper cites Video dif- fusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video dif- fusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.848158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.717471Z digest=sha256:8a8e0c4ece6256f346d398146ce029a62f81f6d457d155f4ef7841ae09b924af

Observation 1b856818-4376-43aa-bc60-d19b3f88d29a · outbound

This paper cites sim- ple diffusion: End-to-end diffusion for high resolution im- ages.

Factorized Video Autoencoders for Efficient Generative Modelling sim- ple diffusion: End-to-end diffusion for high resolution im- ages

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.831083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.723528Z digest=sha256:93b5fe8f335557a5032d209e428f49cf044150efbf0862042a66d864b133adfa

Observation 9362dd53-fa6a-4d1f-9d57-cec8d5f3bce2 · outbound

This paper cites Real-time intermediate flow estimation for video frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Real-time intermediate flow estimation for video frame interpolation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.812209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.728419Z digest=sha256:d3cdbd59d6446fe49989dc0cc2fc68e1222edbb541e8a9efbaada07641e04414

Observation 976a8bc2-a48b-481e-b797-9d52053868cd · outbound

This paper cites Scalable Adaptive Computation for Iterative Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Scalable Adaptive Computation for Iterative Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.733845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.733845Z digest=sha256:d396a8194161f1f5bdbc96419761c00bae70de3113e4daf827775960902839da

Observation c8ce54be-f3a3-4d1f-8a32-23131d49d0ae · outbound

This paper cites Video inter- polation with diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video inter- polation with diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.795613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.739217Z digest=sha256:48d6bae082ae9d88c1ec5dc212030df31a53810d68ad94b852e314b51deb3f36

Observation 19eada01-7cd1-459b-bd51-37fe7aea1571 · outbound

This paper cites Video interpolation with diffu- sion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video interpolation with diffu- sion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.776767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.743897Z digest=sha256:3d44ef957c4c95e5d3a8a6479369ee64d9fbb5746b8b6fa6bad56474055b7796

Observation ed0cd23e-6710-42dc-96e3-cac0cbe6b52c · outbound

This paper cites Benchmarking Video Frame Interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Benchmarking Video Frame Interpolation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.749078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.749078Z digest=sha256:ea8e0074f0a03a460f0386b9c400d68a9a7e00a8ffc9a4bbfd20b7d9a9d0d89d

Observation 16e0ded2-3661-43db-b8a3-d3bd1ae3457d · outbound

This paper cites Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation.

Factorized Video Autoencoders for Efficient Generative Modelling Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.759875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.754098Z digest=sha256:6d55cd7542add6db6034c43165c8ab362091e7760c550fda963a072891a6853a

Observation d67c5c52-ba39-4c39-a17b-f81b38f57dc1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Factorized Video Autoencoders for Efficient Generative Modelling Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.759220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.759220Z digest=sha256:855ba2eab09451c1cadd26c91210d9142212327689379b5f16ea917b4a6fa766

Observation bcf71a2e-c152-455f-b5f9-b1154d11911f · outbound

This paper cites Semcity: Semantic scene gener- ation with triplane diffusion.

Factorized Video Autoencoders for Efficient Generative Modelling Semcity: Semantic scene gener- ation with triplane diffusion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.742631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.766452Z digest=sha256:4ac249942906889e4ee3410147522ea3c95f0cadd92dff387d0b9b2f37145771

Observation 4d88f02c-cbb5-42d7-93c8-b30f7f7bf6c8 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.771423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.771423Z digest=sha256:199498a4b196322a0f7820ed519e634b9e897c7ff5f605c1439492ef41d1beda

Observation e20f972a-40a2-4087-8ac8-b9f900513c40 · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

Factorized Video Autoencoders for Efficient Generative Modelling WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.776579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.776579Z digest=sha256:1a5a6e08add31d7ca760bbdf691b06c93127d795f7dce4cb0c91b8fada36db41

Observation 80853de7-8dda-4104-adcf-8d6a7741b96a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Factorized Video Autoencoders for Efficient Generative Modelling Open-Sora Plan: Open-Source Large Video Generation Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.782354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.782354Z digest=sha256:c8cee6ac31d134455eb188630aedaaf7824ae47b442a7f7a91e6cdbd680958b0

Observation 58eeaaf1-0466-4cbb-8d24-8f5a961ff568 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Factorized Video Autoencoders for Efficient Generative Modelling Common diffusion noise schedules and sample steps are flawed

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.715688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.788060Z digest=sha256:69fadb1f1eee6d39b0cfc6301d10dfde7ffee7d3382701e0734212750b1f0cdd

Observation 89dd405f-52e9-498f-a10b-eb176716b5ad · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.793892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.793892Z digest=sha256:f6d6cb543f9ddf410a66716380662a2adf4ffef2513af151e6babb0715032727

Observation 3b70a688-6960-43da-928f-3946a17155e5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Factorized Video Autoencoders for Efficient Generative Modelling Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.799772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.799772Z digest=sha256:c2428e60248d605ed7342c274a574f58f444264a37f54355f0ea37bd7d98174e

Observation c779d7a1-778d-4067-961f-066afe942f17 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Factorized Video Autoencoders for Efficient Generative Modelling Gener- ating diverse high-fidelity images with vq-vae-2

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.806000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.806000Z digest=sha256:c34aecadc3d864bbe4c154c083eaa4bc073767ed4ab9bd111466742402577e83

Observation 0ee18d90-2f2a-4e4d-8eab-ce93d9dc9c78 · outbound

This paper cites Film: Frame inter- polation for large motion.

Factorized Video Autoencoders for Efficient Generative Modelling Film: Frame inter- polation for large motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.688919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.811064Z digest=sha256:8401ea4ef93fe7bf8b5852d961aa2953cab033469f5bcc5507c0b9c48d7e5215

Observation 0652cc0b-39ec-47ce-9408-5143ba09a015 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021, 2021.

Factorized Video Autoencoders for Efficient Generative Modelling High-resolution image syn- thesis with latent diffusion models, 2021, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.673254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.815511Z digest=sha256:05c08a64430135ff41d4ee599becd84a782c6704142ac4f8f26766d081c5bfef

Observation 39daf284-1542-433c-bb7e-1efee45d9b3d · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Factorized Video Autoencoders for Efficient Generative Modelling U- net: Convolutional networks for biomedical image segmen- tation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.821864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.821864Z digest=sha256:bf306a968043ed57e9591c669aa60695db704cdefaae2b4ead6eb8d857ee7014

Observation dd149793-971a-4dd7-bb34-f8035c81d2b3 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic text-to-image diffusion models with deep language understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.827317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.827317Z digest=sha256:fb71c27a340ba76beac1429b2efcfd09109ae65220f71b5a01793a4a5caa74e2

Observation c24de059-c635-4dea-9e8e-938683fe7d93 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Progressive Distillation for Fast Sampling of Diffusion Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.833018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.833018Z digest=sha256:a77b81e7a5c81022457ee9a74e6daa1a3d46b2de00e1e4a67e48258925349e95

Observation fc48e18a-a25a-4b4c-8049-058f110e5773 · outbound

This paper cites Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein.

Factorized Video Autoencoders for Efficient Generative Modelling Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.635684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.837917Z digest=sha256:7e666e7e1f565a76c2e7ec21b758572fd04c817b12b44b92697a7d9aef58dcf4

Observation 01088514-0af1-4035-bf64-2e370a04a4b8 · outbound

This paper cites Denoising Diffusion Implicit Models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising Diffusion Implicit Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.843325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.843325Z digest=sha256:0c69da809111e84742ff5956017180857f3e5f69becd8df6e697f319beb0ef16

Observation 172af5ce-4c50-4693-a9c7-856b48ed7dfd · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Factorized Video Autoencoders for Efficient Generative Modelling UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.848635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.848635Z digest=sha256:b37a285fac82173733bcf0db11058deff1a6a0ae66e51592f2bbe51747aea823

Observation d46bc2f2-ea24-488c-9762-519615c0e090 · outbound

This paper cites Fvd: A new metric for video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Fvd: A new metric for video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.854150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.854150Z digest=sha256:aac27f2dffbb8898934205004b6e8b642d1ea77917107ea36128d04471f1894b

Observation d4ebc61a-f597-4bf0-aa69-05b2246b6621 · outbound

This paper cites Neural discrete representation learning.

Factorized Video Autoencoders for Efficient Generative Modelling Neural discrete representation learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.859862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.859862Z digest=sha256:2189df9aa49a6a06ace3cc40633e5934f862a49b44cdda9015854cffe22d3a95

Observation 52817f95-b4a4-445d-a685-505eb40dbac1 · outbound

This paper cites Attention is all you need.

Factorized Video Autoencoders for Efficient Generative Modelling Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.865128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.865128Z digest=sha256:4196baf159c435ed8d4f9ebc21c83e0048c2bb08e8e32f5110a4195e132b4d32

Observation 9b1af98f-a6f0-4a51-867e-7489a0ae56f9 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Factorized Video Autoencoders for Efficient Generative Modelling Phenaki: Variable length video generation from open domain textual descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.586719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.870267Z digest=sha256:d501c4e06abb320a6c56ba34244515f5800d0cd5397e9a6c3c30c97b655b338b

Observation f0dbc5c7-6fb3-43ce-a36e-a3d2cadf78d2 · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.568768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.874943Z digest=sha256:1522b8de53822e49b12d87800eccb5a5053821af53633df2209ab15f7b602952

Observation 7f20194b-eaa8-449b-a935-6ab099d44eea · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.880228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.880228Z digest=sha256:65c17ac1f007f03e22f597f6b74d0f242c19c4a3a6830299baf82ba555cc5572

Observation acfc32a9-a8e4-4c21-9880-38064b6f6149 · outbound

This paper cites Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape.

Factorized Video Autoencoders for Efficient Generative Modelling Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.885358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.885358Z digest=sha256:4c3a269d656684be575454a23d0107972469b2e692380dcaa3ca5f24c8e930ca

Observation f0648488-85b3-4d10-8001-0b30b4e4dd01 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Factorized Video Autoencoders for Efficient Generative Modelling CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.892204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.892204Z digest=sha256:379bae04d9d3a6c07476a4409c1533a3196a19635f475037362b5ed17547c640

Observation b7a1ad7b-ef8a-4421-beca-78e70988da9a · outbound

This paper cites Magvit: Masked generative video transformer.

Factorized Video Autoencoders for Efficient Generative Modelling Magvit: Masked generative video transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.539734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.897772Z digest=sha256:4ce9631c68b4eb85205b38385ddc6bb28456a7cd1dac69d0cacb6fc057ec62a4

Observation 3e09ef54-2a39-4041-b5c2-ee92ba68c1e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.903667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.903667Z digest=sha256:585bb15172cb6b5842c8e5e9e247f8078fa77ae56b061baaf6e041e8110b4f74

Observation ef971c59-9385-43bd-8e89-e9c6f5e0118c · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language model beats diffusion - tokenizer is key to visual generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.522805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.910211Z digest=sha256:35d8603cb1634efa17c309bbd0f492b6fd9a5b6abdfaaefd776c40ddb64a7a21

Observation a636f1e9-ef7f-48a5-8d75-9d10fbf6c495 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.915062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.915062Z digest=sha256:5dc4f3280fb2eb9be910d010e2b6c88a26935174aa1ea4d360f31eefd5f9b933

Observation 221e7fb3-0296-4240-b001-28019e484271 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.493940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.919956Z digest=sha256:06d940ed511ffb1a2bc48e4de9930d745a5a5b238437eb064d40035e7983b739

Observation d4d32f16-ae4c-44cc-9bc6-61362294856b · outbound

This paper cites Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition.

Factorized Video Autoencoders for Efficient Generative Modelling Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.925490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.925490Z digest=sha256:3c950898c033d1e75e818f593771ae878a26305de0b0654615db31379db97ae4

Observation b23fa786-5c1f-4aaa-8aa9-facd31e6f960 · outbound

This paper cites Cv- vae: A compatible video vae for latent generative video mod- els.

Factorized Video Autoencoders for Efficient Generative Modelling Cv- vae: A compatible video vae for latent generative video mod- els

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.477573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.930974Z digest=sha256:4cbe9c5c329a63994b63a866bc1f738c7fdc51879490a5840bc27484e4aa52f3

Observation 9ef45027-92b3-45df-a3e2-1f450dd6999e · outbound

This paper cites an unresolved cited work.

Factorized Video Autoencoders for Efficient Generative Modelling Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:30:04.442723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.942929Z digest=sha256:d57a6451ebc46dbfcf1554b18fd4fd2b05b3548af5dbb1b91c93e69a24a152f7

Observation e6b022e0-87ef-404b-b8cc-f93dc781bac3 · outbound

This paper cites For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of.

Factorized Video Autoencoders for Efficient Generative Modelling For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.460419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T21:30:03.937509Z digest=sha256:8481a772fe853bb79c79d7e0badbed867028cca41c4dabb6515f450c08f1dc3c

Observation 25b2559b-ba51-4350-9f14-9132c51ec185 · outbound

This paper cites A Short Note about Kinetics-600.

Factorized Video Autoencoders for Efficient Generative Modelling A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.630717Z digest=sha256:0cad2f9e55e38c98c099b2912cfe38bdd6886062f9b57fc6234cfcadb958e611

Pith citing papers

No inbound Pith citation observations are available.