Pith. sign in

Paper Citation Record · LEDGER

VidTwin: Video VAE with Decoupled Structure and Dynamics

As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2412.17726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17726 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:15.177498Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.523981Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 8a51b4b6-f82f-4d7f-b317-1a0eb6b7b637 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Lumiere: A space-time diffusion model for video generation, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.853191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.833300Z digest=sha256:d25a1f5ed940bf210510b394877040edb40eed2226132aa4e0d29d74b9c9b75b

Observation 30d6ea7a-aead-4b23-a065-0511bdb92929 · outbound

This paper cites Is space-time attention all you need for video understanding?,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Is space-time attention all you need for video understanding?,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.837783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.837783Z digest=sha256:6c1a52e8e2da8b0b2b0626bf7fa4d8ba6dce6faf7d743697f571dcf363956300

Observation 939985a1-b585-43e3-9ab6-fefbafdc4aef · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.841446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.841446Z digest=sha256:0af92b43fe0c00c23925f3e6fb355f305d5244c584cec09c2250cff6e7527f90

Observation d2eef4d4-36a9-40b1-a6ea-a63e82f60e8a · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.828952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.845203Z digest=sha256:d186917b5de83b53e7db7959e3316ded3c4277dbe17fc294c596222c1cda3967

Observation 980c2f6a-7b27-4840-ab6d-b60eaa290259 · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.849034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.849034Z digest=sha256:a370244544ee10b4486026780c452e389af8573b5319cd1cee428e50419ea7c3

Observation cb0d2dec-3db4-4df9-b48c-670b8462b578 · outbound

This paper cites Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.809393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.852749Z digest=sha256:6929ed33903477b57e4c841429af68b8ddb4d7b8db87db539d43d71cbd80b079

Observation 750ba4ef-76fb-4dab-812e-639e7d5dcd1c · outbound

This paper cites Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.797126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.856343Z digest=sha256:073f79ea9ad5c585724f3ce4d25bb86637e1829bbbfe2e0acbcff49aa5bf06d7

Observation 4d5d1d0a-8ddd-4672-975a-c0a0b3bbf3d6 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

VidTwin: Video VAE with Decoupled Structure and Dynamics Taming transformers for high-resolution image synthesis, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.781948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.860340Z digest=sha256:ee3a6133be461bf52c89417e1b926a621ba6387d57c248d60ddfc5c1994d3c0a

Observation d699f7fd-9bbd-415c-958b-8c504c887a79 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.863936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.863936Z digest=sha256:4692fa08e20715368ebf526aa6966fac3266cef636acbe34641809aac41e04c5

Observation efee463a-179c-40b3-9e4b-37ce0ec7c838 · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.762397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.867316Z digest=sha256:211cf5ef8c9ee6dd0417a4a841472603291110c560ccb5774dc537e65369558e

Observation 13b4bfdb-4174-4c7c-9044-17e10ffb8402 · outbound

This paper cites Maskvit: Masked visual pre-training for video prediction, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Maskvit: Masked visual pre-training for video prediction, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.749581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.871468Z digest=sha256:8b0da6501ee5697450014ff5038108b90d2b01244e937ba0c52ac7a2ab1e6f07

Observation 22e6ec2a-5d76-49bf-bcb7-d0e3395697c4 · outbound

This paper cites Gaia: Zero- shot talking avatar generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Gaia: Zero- shot talking avatar generation, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.737987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.874515Z digest=sha256:9c9399faaa13fad6280c0da08db78cf47eef5e04d9457da639086e8d75033791

Observation 6a815379-8151-4e1c-b1ff-8546141cf11b · outbound

This paper cites Latent video diffusion models for high-fidelity long video generation, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Latent video diffusion models for high-fidelity long video generation, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.726603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.877971Z digest=sha256:9f33be74bd3e58b8fdea51eaed131fb95b38db69cbb5f19b8ad9a999bfe44a46

Observation cc329618-d09e-4155-b6d9-3801ce893388 · outbound

This paper cites Classifier-free diffusion guidance, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Classifier-free diffusion guidance, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.715863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.881304Z digest=sha256:09dd24621d732ab2fb28edd1d8356f51221902ce66800ea8fcdbe6c6d275b48f

Observation 9a5eca5c-96dc-49aa-8ed2-5b40b5dbfbb3 · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

VidTwin: Video VAE with Decoupled Structure and Dynamics Denoising diffu- sion probabilistic models, 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.704214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.884748Z digest=sha256:a203bb3eb81db6ee59ff699c6c5955df6f505ac108725174e8e2b1bda8bda8e5

Observation 5817e66a-784b-44d6-a5b1-96bb176da9db · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.691580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.888244Z digest=sha256:b5a8c5da025552de7bd7b5abedfd568c874e708deab4b9626a61d834453bbca6

Observation 3bdea22d-7a28-4964-ab87-2f792ddb7244 · outbound

This paper cites an unresolved cited work.

VidTwin: Video VAE with Decoupled Structure and Dynamics Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:20:15.677462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.892694Z digest=sha256:994c8b07179dd2016484b276711bbaa947d0a994e8d2d69eeb13475d2addb161

Observation b7084fb8-1a02-47d7-8ede-a0f0cea68e7f · outbound

This paper cites Image quality metrics: Psnr vs.

VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality metrics: Psnr vs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.896687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.896687Z digest=sha256:3bc00c82c76dc143bba0fbdfaacc95c6c7f3d73d229328d7b49dc59cc5e022a5

Observation a4262265-11ba-48d2-82e6-473c74470020 · outbound

This paper cites Dive: Dit-based video generation with enhanced control, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Dive: Dit-based video generation with enhanced control, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.659864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.900677Z digest=sha256:4467a61ebdc1745249f58a485a021a973c7aa286f159ce6e2b519908c57d5aa1

Observation cb2d66d6-ce0b-47b0-97a8-8a025b5a5259 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.648763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.904194Z digest=sha256:53198f450dc7b10381ad4b26d70e8673aca9edd0735d3a7488b4011c2dabf7b9

Observation 423a01fb-3249-47f8-bdad-d8586f9bae21 · outbound

This paper cites Kingma and Jimmy Ba.

VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma and Jimmy Ba

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.636751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.908332Z digest=sha256:f791151419f87472489248f19a7818e318bcc784a12d9e8742fad6c51d70b48e

Observation 455d1cb4-25aa-493a-b752-b233614df139 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Auto-encoding varia- tional bayes, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.626363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.911730Z digest=sha256:8bdc4d61081b93044d2901290c020a04d917804a50fe92560c8b14e58a8fd042

Observation 8f0fb8fa-0c4a-43da-903d-3a7751493cea · outbound

This paper cites Mpeg: A video compression standard for multimedia applications.

VidTwin: Video VAE with Decoupled Structure and Dynamics Mpeg: A video compression standard for multimedia applications

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.614945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.914950Z digest=sha256:a8fdc222ab09fc79fb62106c2ae45e74f41da25b499b540b0e6102dfaefe8188

Observation 9de07438-a2cb-4fb5-b457-f70e028a892c · outbound

This paper cites Disentangled motion modeling for video frame interpolation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Disentangled motion modeling for video frame interpolation, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.603104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.918319Z digest=sha256:3bd4d75cd434105f01f8b39b33d923c36042c25668ca690376df8e15793c673f

Observation 052ce252-8499-4549-a593-102c8f8c9028 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.591056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.921385Z digest=sha256:2010873388795c9c86ab53eb8dd0acb1f4e65618427730a24f3c71e2e0c2fadb

Observation 924c9bcf-4a3a-45d8-a213-923bbb3b3393 · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.578327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.924690Z digest=sha256:4fbefa4bc2256d633da79ecb95d29395c5f5223ca6b71f4579ad1ef87f4b22c9

Observation bd3a3ad6-4eda-4771-9c63-771799d62655 · outbound

This paper cites Open-sora plan: Open-source large video generation model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora plan: Open-source large video generation model, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.567787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.928859Z digest=sha256:6bfb092bb1f698d9d8ce4ebf210c32e130887def8f72ed38b06517fa5bd5c86b

Observation 87ad3afc-bb50-4f2d-b28f-78a0058a8530 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Finite scalar quantization: Vq-vae made simple, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.556722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.932598Z digest=sha256:01ac07bea0d964922bd6f7b2a39cec30c597601302454cd452fa6e4c15380c69

Observation e510d70c-295b-438a-b4cd-be509f36a39e · outbound

This paper cites Video generation models as world simulators.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video generation models as world simulators

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.545131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.936454Z digest=sha256:8cd8dab2e1903fd884a2319590fc7c3c0588524976fc14f2795d9daa6ac7336d

Observation 173455d7-a39f-4480-b391-802264b326a1 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Scalable diffusion models with transformers, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.532675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.940095Z digest=sha256:f6be7c34c092a503d94496f6555e865c0b887f57c26b4cef79125f5f03cf682e

Observation ad9038f8-7450-447d-81e5-5dc39d72f040 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.943702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.943702Z digest=sha256:f6979ba46571add797857b85c46b7cf36bb2cc7cae9c64c24eb76dfc4eac1c63

Observation 6643c645-fc39-491d-8589-c90136903690 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VidTwin: Video VAE with Decoupled Structure and Dynamics High-resolution image synthesis with latent diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.947305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.947305Z digest=sha256:6094fe35f78ba90da8d6cca6a6fc786215a9ccd375c2d2bcbad5a57170401278

Observation 37570def-13f5-445f-892b-59827a87c91f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

VidTwin: Video VAE with Decoupled Structure and Dynamics U-net: Convolutional networks for biomedical image segmentation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.950692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.950692Z digest=sha256:c3c3ffecaa2e6efbacf60e094baf27b01c441263550065cc2babb92ed959ac5e

Observation 51ee7803-ea9c-44c0-abdc-a0cd568b579f · outbound

This paper cites Adversarial diffusion distillation, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Adversarial diffusion distillation, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.503772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.954196Z digest=sha256:5c53f62fde6b9b90d519a8fe91263df52e6334d0fb8f24b05a04206413f293a3

Observation fe8ac751-be3d-49fb-970b-7726cbe6aedc · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.957686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.957686Z digest=sha256:0d18dee538d81231324174d5e961064d07195cc2937b9f841f42f2e821ea1068

Observation 97e5951a-f6f9-429e-b587-3d0f9f4bab8a · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Make-a-video: Text-to-video generation without text-video data, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.485294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.961306Z digest=sha256:3e2954092a9d18b8ba02addbbf86895a3a6f4deb8988f10e6f99dd9d658509cc

Observation 5580f9b1-ae07-4cc5-a575-7fd1d4940c33 · outbound

This paper cites Denois- ing diffusion implicit models, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Denois- ing diffusion implicit models, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.472440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.965558Z digest=sha256:cb1c224d328f6b1f2a520e1fb2347a4efc86d3cb9fc896eec7c171db04f019cc

Observation d6b1eed6-8857-47d4-a42f-bab53e265fea · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

VidTwin: Video VAE with Decoupled Structure and Dynamics Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.460313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.968894Z digest=sha256:e5944240f280dae8851767de8fa6c8e6757a5e4b29c389a81e33b96a2ff92bf3

Observation 85fa5572-3bdd-41d3-ba03-3cfb56585fa4 · outbound

This paper cites VidTok: A Versatile and Open-Source Video Tokenizer.

VidTwin: Video VAE with Decoupled Structure and Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.972158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.972158Z digest=sha256:6b8eb0c290f5f05628a42f6872bf547fc493676bec6a4308fa56f48a4177cc09

Observation 19ea52da-1ee8-453c-b874-64c470e31fc4 · outbound

This paper cites To- wards accurate generative models of video: A new metric & challenges, 2019.

VidTwin: Video VAE with Decoupled Structure and Dynamics To- wards accurate generative models of video: A new metric & challenges, 2019

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.976149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.976149Z digest=sha256:ca012dc6097032ed84b9cae336a0ad8db2ed40055847ad7a89d5bce2db7b07e0

Observation d9876e47-d21f-4e97-a5f6-1c5cacd4a381 · outbound

This paper cites Neural discrete representation learning,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Neural discrete representation learning,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.979847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.979847Z digest=sha256:6b04013069bc46f6ea93c0ba216db16c7c3c1041bf7d40b7ee8ca5941731e832

Observation d987cf8c-d8e6-4114-9d4d-f3c0598db836 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

VidTwin: Video VAE with Decoupled Structure and Dynamics Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.435576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.983668Z digest=sha256:1a9832720fda1a93028f412e57cb643a5d1929738dddf9f20bb165c406d27812

Observation ca92a608-5a65-4c86-9430-df41fdc0d127 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual description, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Phenaki: Variable length video generation from open domain textual description, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.421657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.986907Z digest=sha256:a738a521f46b9a52b140aecc8c87e86f49be5cebcb418f68809903a2b82a92c7

Observation 807e9664-9048-4004-b4f8-8fcc69693e8c · outbound

This paper cites Mcl-jcv: a jnd-based h.

VidTwin: Video VAE with Decoupled Structure and Dynamics Mcl-jcv: a jnd-based h

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.410180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.990360Z digest=sha256:9370888e89f500ab135b162c1999e73cd722dd2ba131418809c00f4d45cb5ba6

Observation 7de20dd2-15e5-4a32-8ee6-42a7acf73a27 · outbound

This paper cites Bevt: Bert pretraining of video transformers,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Bevt: Bert pretraining of video transformers,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.398791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.994040Z digest=sha256:1d629308cf50e94e4a0899ec23f86de63e1f3a59d77de78147ede72bd713b271

Observation 9f6493ea-3575-49f0-a9dd-a49a5b66912f · outbound

This paper cites Emu3: Next-token prediction is all you need, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Emu3: Next-token prediction is all you need, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.387109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:14.997845Z digest=sha256:7fd7df9cf7a8902a2598a6946774895f8650d6ffec1d2bb81f46d7fda72c7eae

Observation 6290c8bc-e4ca-4e51-b91b-6673e02eda3e · outbound

This paper cites Instructavatar: Text- guided emotion and motion control for avatar generation,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Instructavatar: Text- guided emotion and motion control for avatar generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.376176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.002566Z digest=sha256:3ef6ec65e82fcfe4d4adb7832a89ca462dfcff6fdf03a1cdaac834df1c17b1b9

Observation 795e550f-68be-4473-b774-9977d0adb0a3 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.007189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.007189Z digest=sha256:92223a1af2e990eb2ff9965032c272a176f9035ba540c076e816836b2db57fc2

Observation 74505f1e-79ca-43f0-bd80-31a878d1698e · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality assessment: from error visibility to structural similarity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.012795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.012795Z digest=sha256:f99311487a094c135785105b5915094fa5b043b605fb76b3ef343f47b93df5f7

Observation 5c843c61-c080-42dc-86c6-f7357543a426 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.354098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.018731Z digest=sha256:52afb94a94d6104817e2100a1562f5d19d677c13421d630f41637163dbee953d

Observation 03644ed3-7253-4b81-aa0c-fc0cb8e1b16e · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics ivideogpt: Interactive videogpts are scalable world models, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.343791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.023218Z digest=sha256:e570b68ed81104d7f82d7bcb6d207ae815bb1948050ce8842489c38e68dae644

Observation 1c4b000f-6b2b-4dc8-8f70-c2d6a6fe5f46 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

VidTwin: Video VAE with Decoupled Structure and Dynamics Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.332469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.027510Z digest=sha256:e004df7cebe7d94505f5d65c0329bd69868c145e259822c5315d10e6780b5003

Observation 8fe3a8d6-92aa-4ce5-963c-5cbf606de626 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Cogvideox: Text-to-video diffusion models with an expert transformer, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.322537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.031917Z digest=sha256:eb484f628df488674200e62e7af592c02392935b2d48bd7dd55f5c3aa9a30bf6

Observation dda1ea4e-8b0b-4d2f-a6fe-193a372632fc · outbound

This paper cites Vector-quantized image modeling with improved vqgan, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Vector-quantized image modeling with improved vqgan, 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.036343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.036343Z digest=sha256:d8842e1eb1e587a03b95756351065aef5d99013d97b2024418db40de9b375822

Observation 7a3dfeba-01ee-499e-b89e-8fd5b3249106 · outbound

This paper cites Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

VidTwin: Video VAE with Decoupled Structure and Dynamics Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.303409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.040599Z digest=sha256:6c81d511afb35d027b8786840af7159961137f1390532876fd619002a56ceec6

Observation 4acacae4-5516-4dec-8315-c1db71da2988 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

VidTwin: Video VAE with Decoupled Structure and Dynamics Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.044896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.044896Z digest=sha256:3c21faa3122539856579b9e4f3fce7c7f8b7544a424eea34286d9f7c62ed3ac9

Observation d1bcb5cb-a46b-4c0b-9509-f74b5073fe01 · outbound

This paper cites Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.291433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.049116Z digest=sha256:a91b81bf2040f954e292561e5f1c62c3c8e69b524652d791b8b6b9551cf67ade

Observation 1347ab3e-fe86-4e52-8d7b-cd36406f7d22 · outbound

This paper cites Efficient video diffusion models via content-frame motion-latent decomposition, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Efficient video diffusion models via content-frame motion-latent decomposition, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.278532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.053058Z digest=sha256:c76d0da6f28b2a48c2f61e1ff398b0f1d0f5051a9f4ae9ec2094c3808c8bdd8a

Observation dfd38271-7304-4290-8088-3d54a0da7dda · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

VidTwin: Video VAE with Decoupled Structure and Dynamics Efros, Eli Shecht- man, and Oliver Wang

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.056332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.056332Z digest=sha256:92c8aa2ae9ef21431ce6a779d4cefab6315d7821cd62be619d3f6d1ef56ea066

Observation e66641f7-ee8c-470f-b8f1-6105676630a9 · outbound

This paper cites Video in-context learning: Autore- gressive transformers are zero-shot video imitators.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video in-context learning: Autore- gressive transformers are zero-shot video imitators

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.261125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.059760Z digest=sha256:3de276475f519f505b40fb39c8a270a3c1e41b7a9a6d38e727d24e761f63cee1

Observation c4fa64eb-468e-42b6-ad53-0923e4ad0126 · outbound

This paper cites Cv-vae: A compatible video vae for latent generative video models,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Cv-vae: A compatible video vae for latent generative video models,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.248414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.173588Z digest=sha256:f67077b5f4274aed26c3e0fcbcfab866939248cff5b0a5b40427aa83a22376af

Observation b91d5f28-1d4c-458a-96e9-693274ea7182 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora: Democratizing efficient video production for all, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.237026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T05:20:15.177498Z digest=sha256:cf6a1f679c73201c9e1755a5289ab996ca89f19dd9aa958c1eb28a0b222ac90f

Pith citing papers

Observation abf58d5f-9249-463d-8ac4-8bed29cccf68 · inbound

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion cites this paper.

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:17.395406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:17.395406Z digest=sha256:7ad37f62bd6b0fa32fe3d3eec73922412fb6a1b6d343e0a876dc0a4cea695922

Observation ef828168-849f-4259-8d45-f7f2490571e5 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:a3e803a3ff02b369976b6c064e113be6678f31be3d830d7a35cd4e2964150d00

Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · inbound

V-RAE: Rethinking Video Latent Spaces for Generation cites this paper.

V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.523981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.523981Z digest=sha256:60c8522344268d3c6e71dc83e7aace73648020355d4d79ad20f396b055d46a29