Pith. sign in

Paper Citation Record · LEDGER

VidTwin: Video VAE with Decoupled Structure and Dynamics

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2412.17726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17726 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:15.177498Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.523981Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 8a51b4b6-f82f-4d7f-b317-1a0eb6b7b637 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Lumiere: A space-time diffusion model for video generation, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.853191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.833300Z digest=sha256:f4cb0169d003c066929385dcabba18e7a71295a2e222656a67586b76955c48d8

Observation 30d6ea7a-aead-4b23-a065-0511bdb92929 · outbound

This paper cites Is space-time attention all you need for video understanding?,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Is space-time attention all you need for video understanding?,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.837783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.837783Z digest=sha256:6c1a52e8e2da8b0b2b0626bf7fa4d8ba6dce6faf7d743697f571dcf363956300

Observation 939985a1-b585-43e3-9ab6-fefbafdc4aef · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.841446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.841446Z digest=sha256:0af92b43fe0c00c23925f3e6fb355f305d5244c584cec09c2250cff6e7527f90

Observation d2eef4d4-36a9-40b1-a6ea-a63e82f60e8a · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.828952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.845203Z digest=sha256:df919cefe28c8d124f622ce08c54c5f877ce98f746e9b6cdb653364a932404ef

Observation 980c2f6a-7b27-4840-ab6d-b60eaa290259 · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.849034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.849034Z digest=sha256:a370244544ee10b4486026780c452e389af8573b5319cd1cee428e50419ea7c3

Observation cb0d2dec-3db4-4df9-b48c-670b8462b578 · outbound

This paper cites Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.809393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.852749Z digest=sha256:b168e20be109797a085d311444c4c8cdbe12060d970b12f75e3a8a06534c5162

Observation 750ba4ef-76fb-4dab-812e-639e7d5dcd1c · outbound

This paper cites Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.797126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.856343Z digest=sha256:53130e26a1f6a10aed2d3a8d8012ae84272b97d4bb8edfe528209b8783cd30ac

Observation 4d5d1d0a-8ddd-4672-975a-c0a0b3bbf3d6 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

VidTwin: Video VAE with Decoupled Structure and Dynamics Taming transformers for high-resolution image synthesis, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.781948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.860340Z digest=sha256:36509a2678c1bb0592c1f11df659bd95f424365912584f2f33e111e7068d4106

Observation d699f7fd-9bbd-415c-958b-8c504c887a79 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Scaling rectified flow trans- formers for high-resolution image synthesis, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.863936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.863936Z digest=sha256:4692fa08e20715368ebf526aa6966fac3266cef636acbe34641809aac41e04c5

Observation efee463a-179c-40b3-9e4b-37ce0ec7c838 · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.762397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.867316Z digest=sha256:124ab3d7f959e363fc0c97632b8459381c74aab82edb92d0324fb803deb22bd9

Observation 13b4bfdb-4174-4c7c-9044-17e10ffb8402 · outbound

This paper cites Maskvit: Masked visual pre-training for video prediction, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Maskvit: Masked visual pre-training for video prediction, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.749581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.871468Z digest=sha256:c38740986839da713aeece142582ce354a76a8099ce8146e0392d1768b3fba2c

Observation 22e6ec2a-5d76-49bf-bcb7-d0e3395697c4 · outbound

This paper cites Gaia: Zero- shot talking avatar generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Gaia: Zero- shot talking avatar generation, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.737987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.874515Z digest=sha256:db2bbc480ca93a7dfb86d677a995414fe456e7454d0c6ecd2f289746a74a796e

Observation 6a815379-8151-4e1c-b1ff-8546141cf11b · outbound

This paper cites Latent video diffusion models for high-fidelity long video generation, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Latent video diffusion models for high-fidelity long video generation, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.726603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.877971Z digest=sha256:a12892fa358d6f8abb66ae73f6997c9e6293f7c4c1bf7c90a8fee1905ecd77ff

Observation cc329618-d09e-4155-b6d9-3801ce893388 · outbound

This paper cites Classifier-free diffusion guidance, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Classifier-free diffusion guidance, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.715863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.881304Z digest=sha256:2f971c05b449d08a91ca2daeff9e59ac19ec8cd02e0c221e12bc98fe7081dab0

Observation 9a5eca5c-96dc-49aa-8ed2-5b40b5dbfbb3 · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

VidTwin: Video VAE with Decoupled Structure and Dynamics Denoising diffu- sion probabilistic models, 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.704214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.884748Z digest=sha256:fd76df3896ae1d088c3928576f7030f4435c3d267eaa427429da375dbed25f38

Observation 5817e66a-784b-44d6-a5b1-96bb176da9db · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.691580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.888244Z digest=sha256:6a12908da2bc9589ef95eb7010ae15504fa77a1174807288907da0d7b2f11520

Observation 3bdea22d-7a28-4964-ab87-2f792ddb7244 · outbound

This paper cites an unresolved cited work.

VidTwin: Video VAE with Decoupled Structure and Dynamics Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:20:15.677462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.892694Z digest=sha256:f3c62034c3d404e21a1dc0e3bb6c8753d1d7edf356092b60302b4bca117e9180

Observation b7084fb8-1a02-47d7-8ede-a0f0cea68e7f · outbound

This paper cites Image quality metrics: Psnr vs.

VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality metrics: Psnr vs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.896687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.896687Z digest=sha256:3bc00c82c76dc143bba0fbdfaacc95c6c7f3d73d229328d7b49dc59cc5e022a5

Observation a4262265-11ba-48d2-82e6-473c74470020 · outbound

This paper cites Dive: Dit-based video generation with enhanced control, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Dive: Dit-based video generation with enhanced control, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.659864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.900677Z digest=sha256:6956cbfd6c0afbfb02f44a60b10f6dbe869de8f43cb6c7186553414ea8358c2a

Observation cb2d66d6-ce0b-47b0-97a8-8a025b5a5259 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.648763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.904194Z digest=sha256:d55b3c9b0b6149814927dd34fe5a9bd27fce2ae1242d059a5d57caa6f4bc6094

Observation 423a01fb-3249-47f8-bdad-d8586f9bae21 · outbound

This paper cites Kingma and Jimmy Ba.

VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma and Jimmy Ba

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.636751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.908332Z digest=sha256:a3d4db7db074b0a09cfd54f2da5edec4d8ba0ce98a0648d8ebc7994dc63b2ba6

Observation 455d1cb4-25aa-493a-b752-b233614df139 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Auto-encoding varia- tional bayes, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.626363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.911730Z digest=sha256:60068417eb83d212d7b634dc22470240c835cb56297afce4a9ea2364c0d11eeb

Observation 8f0fb8fa-0c4a-43da-903d-3a7751493cea · outbound

This paper cites Mpeg: A video compression standard for multimedia applications.

VidTwin: Video VAE with Decoupled Structure and Dynamics Mpeg: A video compression standard for multimedia applications

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.614945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.914950Z digest=sha256:e1473e49684e6fb25c48e36bf4a1803b17ba77aef7bca2f0d72b68d946e4a23a

Observation 9de07438-a2cb-4fb5-b457-f70e028a892c · outbound

This paper cites Disentangled motion modeling for video frame interpolation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Disentangled motion modeling for video frame interpolation, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.603104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.918319Z digest=sha256:3ba0c8edb46b1c0f446031144683d7fe6125b04ebc62e439350db46e8fb8248c

Observation 052ce252-8499-4549-a593-102c8f8c9028 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.591056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.921385Z digest=sha256:49104f34f4ca7faa90b2e33aad7698c2a1a81db8fad61e039cf1fae719bd03bb

Observation 924c9bcf-4a3a-45d8-a213-923bbb3b3393 · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.578327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.924690Z digest=sha256:1f018dba1dc4465bc3f47a8e6ea8d4b58c56510a186b5122653d33fea2b196e6

Observation bd3a3ad6-4eda-4771-9c63-771799d62655 · outbound

This paper cites Open-sora plan: Open-source large video generation model, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora plan: Open-source large video generation model, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.567787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.928859Z digest=sha256:f9136b30943676826d6c801ea80daf25dbb10c07e38f05fa63ce913c2cf482c3

Observation 87ad3afc-bb50-4f2d-b28f-78a0058a8530 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Finite scalar quantization: Vq-vae made simple, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.556722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.932598Z digest=sha256:b8489b85923914478411630d695127ddc9e0921a9f4d051b799be5ce0d075f5f

Observation e510d70c-295b-438a-b4cd-be509f36a39e · outbound

This paper cites Video generation models as world simulators.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video generation models as world simulators

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.545131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.936454Z digest=sha256:a0d3dc371294cfa11f09e8a36a961ba16a2abe3caa29af0b37877ac07c600256

Observation 173455d7-a39f-4480-b391-802264b326a1 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Scalable diffusion models with transformers, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.532675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.940095Z digest=sha256:44703600408c2ad88bced7841cf216fea9271e5aa01152f6edfc92bd49f9c71e

Observation ad9038f8-7450-447d-81e5-5dc39d72f040 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.943702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.943702Z digest=sha256:f6979ba46571add797857b85c46b7cf36bb2cc7cae9c64c24eb76dfc4eac1c63

Observation 6643c645-fc39-491d-8589-c90136903690 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VidTwin: Video VAE with Decoupled Structure and Dynamics High-resolution image synthesis with latent diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.947305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.947305Z digest=sha256:6094fe35f78ba90da8d6cca6a6fc786215a9ccd375c2d2bcbad5a57170401278

Observation 37570def-13f5-445f-892b-59827a87c91f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

VidTwin: Video VAE with Decoupled Structure and Dynamics U-net: Convolutional networks for biomedical image segmentation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.950692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.950692Z digest=sha256:c3c3ffecaa2e6efbacf60e094baf27b01c441263550065cc2babb92ed959ac5e

Observation 51ee7803-ea9c-44c0-abdc-a0cd568b579f · outbound

This paper cites Adversarial diffusion distillation, 2023.

VidTwin: Video VAE with Decoupled Structure and Dynamics Adversarial diffusion distillation, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.503772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.954196Z digest=sha256:6230d611de006c5f47ea0a2680ba6918f3485e5089310b22fc81920bad798540

Observation fe8ac751-be3d-49fb-970b-7726cbe6aedc · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.957686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.957686Z digest=sha256:0d18dee538d81231324174d5e961064d07195cc2937b9f841f42f2e821ea1068

Observation 97e5951a-f6f9-429e-b587-3d0f9f4bab8a · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Make-a-video: Text-to-video generation without text-video data, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.485294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.961306Z digest=sha256:e1c770b5845c6eb35400f690011d7afee74f2bb4568a0134a391e1001b7774f9

Observation 5580f9b1-ae07-4cc5-a575-7fd1d4940c33 · outbound

This paper cites Denois- ing diffusion implicit models, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Denois- ing diffusion implicit models, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.472440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.965558Z digest=sha256:aef6a1e648e4d3f63a083118335382eb10ec26eb2040de2912286a2db7f98374

Observation d6b1eed6-8857-47d4-a42f-bab53e265fea · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

VidTwin: Video VAE with Decoupled Structure and Dynamics Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.460313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.968894Z digest=sha256:a034fed0548454a3a5d7cb8956ef53905ba2710292913840ac4657c37675dc59

Observation 85fa5572-3bdd-41d3-ba03-3cfb56585fa4 · outbound

This paper cites VidTok: A Versatile and Open-Source Video Tokenizer.

VidTwin: Video VAE with Decoupled Structure and Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.972158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.972158Z digest=sha256:0d67ed064952197072d01f1009f3ac6010edde4c625c0acb416fa6bc5d138bda

Observation 19ea52da-1ee8-453c-b874-64c470e31fc4 · outbound

This paper cites To- wards accurate generative models of video: A new metric & challenges, 2019.

VidTwin: Video VAE with Decoupled Structure and Dynamics To- wards accurate generative models of video: A new metric & challenges, 2019

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.976149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.976149Z digest=sha256:ca012dc6097032ed84b9cae336a0ad8db2ed40055847ad7a89d5bce2db7b07e0

Observation d9876e47-d21f-4e97-a5f6-1c5cacd4a381 · outbound

This paper cites Neural discrete representation learning,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Neural discrete representation learning,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.979847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.979847Z digest=sha256:6b04013069bc46f6ea93c0ba216db16c7c3c1041bf7d40b7ee8ca5941731e832

Observation d987cf8c-d8e6-4114-9d4d-f3c0598db836 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

VidTwin: Video VAE with Decoupled Structure and Dynamics Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.435576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.983668Z digest=sha256:aebe772faf8ae53ca2f6fe57309f0cb50dc0fbb187fe57f46fe08bb85d32bf3f

Observation ca92a608-5a65-4c86-9430-df41fdc0d127 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual description, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Phenaki: Variable length video generation from open domain textual description, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.421657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.986907Z digest=sha256:d5bc424a4e68ae680fdfb5a4c3125a5c4b1b656323163c8678b101b50ec2ed0c

Observation 807e9664-9048-4004-b4f8-8fcc69693e8c · outbound

This paper cites Mcl-jcv: a jnd-based h.

VidTwin: Video VAE with Decoupled Structure and Dynamics Mcl-jcv: a jnd-based h

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.410180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.990360Z digest=sha256:2dc7d2f42f355e4e29385b56875bf56093bb6d848d2b2be103bbd869fab7238f

Observation 7de20dd2-15e5-4a32-8ee6-42a7acf73a27 · outbound

This paper cites Bevt: Bert pretraining of video transformers,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Bevt: Bert pretraining of video transformers,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.398791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.994040Z digest=sha256:29e1c2dbc33260632b653115a41bb0816501abab559322305c04234950425085

Observation 9f6493ea-3575-49f0-a9dd-a49a5b66912f · outbound

This paper cites Emu3: Next-token prediction is all you need, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Emu3: Next-token prediction is all you need, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.387109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:14.997845Z digest=sha256:bb3e86ec2c11a6086b27c7cb485bc42786bd749c6ac12976c5715d503c541a8d

Observation 6290c8bc-e4ca-4e51-b91b-6673e02eda3e · outbound

This paper cites Instructavatar: Text- guided emotion and motion control for avatar generation,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Instructavatar: Text- guided emotion and motion control for avatar generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.376176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.002566Z digest=sha256:2a8d42101fdba531ea1a8c86505bc5f6cb244ed77c218c8ff335554dffdda9e1

Observation 795e550f-68be-4473-b774-9977d0adb0a3 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.007189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.007189Z digest=sha256:92223a1af2e990eb2ff9965032c272a176f9035ba540c076e816836b2db57fc2

Observation 74505f1e-79ca-43f0-bd80-31a878d1698e · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality assessment: from error visibility to structural similarity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.012795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.012795Z digest=sha256:f99311487a094c135785105b5915094fa5b043b605fb76b3ef343f47b93df5f7

Observation 5c843c61-c080-42dc-86c6-f7357543a426 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.354098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.018731Z digest=sha256:d7c06652714a76ae3051b56d39e52c01ee5bdaebc0cd9c3d8db73fafcaad3aab

Observation 03644ed3-7253-4b81-aa0c-fc0cb8e1b16e · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics ivideogpt: Interactive videogpts are scalable world models, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.343791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.023218Z digest=sha256:116939bdae8692adce103c9951e7ddf2392c756a5a10f4911922e55ab5053cf4

Observation 1c4b000f-6b2b-4dc8-8f70-c2d6a6fe5f46 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

VidTwin: Video VAE with Decoupled Structure and Dynamics Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.332469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.027510Z digest=sha256:0095fe08afb1a6d28ee1eb3793538fef6f7dfd438bcfb19addc345f797985030

Observation 8fe3a8d6-92aa-4ce5-963c-5cbf606de626 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Cogvideox: Text-to-video diffusion models with an expert transformer, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.322537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.031917Z digest=sha256:37e3873ec3d97e0b7052f4174a39c29dc899d29849aa5923e2a1eb70b99fd495

Observation dda1ea4e-8b0b-4d2f-a6fe-193a372632fc · outbound

This paper cites Vector-quantized image modeling with improved vqgan, 2022.

VidTwin: Video VAE with Decoupled Structure and Dynamics Vector-quantized image modeling with improved vqgan, 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.036343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.036343Z digest=sha256:d8842e1eb1e587a03b95756351065aef5d99013d97b2024418db40de9b375822

Observation 7a3dfeba-01ee-499e-b89e-8fd5b3249106 · outbound

This paper cites Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

VidTwin: Video VAE with Decoupled Structure and Dynamics Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.303409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.040599Z digest=sha256:a1b6c57c04f2a203832ec65f9bd2b38d05710e8d8fe69de1df56b6dfcd4b6099

Observation 4acacae4-5516-4dec-8315-c1db71da2988 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

VidTwin: Video VAE with Decoupled Structure and Dynamics Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.044896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.044896Z digest=sha256:3c21faa3122539856579b9e4f3fce7c7f8b7544a424eea34286d9f7c62ed3ac9

Observation d1bcb5cb-a46b-4c0b-9509-f74b5073fe01 · outbound

This paper cites Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.291433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.049116Z digest=sha256:3b2d71b86e3165bd5fc3d93159d00caf3bc82bb08bad293176d17fa304249231

Observation 1347ab3e-fe86-4e52-8d7b-cd36406f7d22 · outbound

This paper cites Efficient video diffusion models via content-frame motion-latent decomposition, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Efficient video diffusion models via content-frame motion-latent decomposition, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.278532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.053058Z digest=sha256:c87d1fed11af53620e64e5e10894f3c22f45018b6900a5970009cbc6d69432d3

Observation dfd38271-7304-4290-8088-3d54a0da7dda · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

VidTwin: Video VAE with Decoupled Structure and Dynamics Efros, Eli Shecht- man, and Oliver Wang

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:15.056332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:15.056332Z digest=sha256:92c8aa2ae9ef21431ce6a779d4cefab6315d7821cd62be619d3f6d1ef56ea066

Observation e66641f7-ee8c-470f-b8f1-6105676630a9 · outbound

This paper cites Video in-context learning: Autore- gressive transformers are zero-shot video imitators.

VidTwin: Video VAE with Decoupled Structure and Dynamics Video in-context learning: Autore- gressive transformers are zero-shot video imitators

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.261125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.059760Z digest=sha256:5607e08c17dde3c65724286fdfdccc1fbbb3904c6c300e972dbdf3a9a1459bd5

Observation c4fa64eb-468e-42b6-ad53-0923e4ad0126 · outbound

This paper cites Cv-vae: A compatible video vae for latent generative video models,.

VidTwin: Video VAE with Decoupled Structure and Dynamics Cv-vae: A compatible video vae for latent generative video models,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.248414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.173588Z digest=sha256:a69db554648321767055457e812a759924125dfcf738ed90a0fbceeaaf1027ac

Observation b91d5f28-1d4c-458a-96e9-693274ea7182 · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora: Democratizing efficient video production for all, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:20:15.237026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T05:20:15.177498Z digest=sha256:1f18253a21f534c08ffbf28549a1eaa1691711537a44270edf367b01c66dc1ef

Pith citing papers

Observation abf58d5f-9249-463d-8ac4-8bed29cccf68 · inbound

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion cites this paper.

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:17.395406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:17.395406Z digest=sha256:ab911642db7ec727d23d93bffc8b8f2db62c148faa4bb528b526d6cb319b40cd

Observation ef828168-849f-4259-8d45-f7f2490571e5 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:be117d64881d953f7e9fba9210634d6bcce4926a1f678b9e3fdd10fb8894f379

Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · inbound

V-RAE: Rethinking Video Latent Spaces for Generation cites this paper.

V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.523981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.523981Z digest=sha256:6dd35d504e6b21343e83cb73f195ead8f773436dce53503edea9e31fc04bd127