Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:15.177498Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 3 inbound Pith citation observations for arXiv:2412.17726.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:15.177498Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.523981Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T05:30:23.456663Z
62 of 62 outbound references displayed
External citation measurements
0
pith, observed 2026-08-10T05:30:23.456663Z
Observation 8a51b4b6-f82f-4d7f-b317-1a0eb6b7b637 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Lumiere: A space-time diffusion model for video generation, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 30d6ea7a-aead-4b23-a065-0511bdb92929 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Is space-time attention all you need for video understanding?,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939985a1-b585-43e3-9ab6-fefbafdc4aef · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eef4d4-36a9-40b1-a6ea-a63e82f60e8a · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 980c2f6a-7b27-4840-ab6d-b60eaa290259 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0d2dec-3db4-4df9-b48c-670b8462b578 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Pixart-sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 750ba4ef-76fb-4dab-812e-639e7d5dcd1c · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Od-vae: An omni-dimensional video compressor for im- proving latent video diffusion model, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4d5d1d0a-8ddd-4672-975a-c0a0b3bbf3d6 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Taming transformers for high-resolution image synthesis, 2021
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d699f7fd-9bbd-415c-958b-8c504c887a79 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Scaling rectified flow trans- formers for high-resolution image synthesis, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efee463a-179c-40b3-9e4b-37ce0ec7c838 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Long video generation with time-agnostic vqgan and time- sensitive transformer, 2022
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13b4bfdb-4174-4c7c-9044-17e10ffb8402 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Maskvit: Masked visual pre-training for video prediction, 2022
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 22e6ec2a-5d76-49bf-bcb7-d0e3395697c4 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Gaia: Zero- shot talking avatar generation, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a815379-8151-4e1c-b1ff-8546141cf11b · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Latent video diffusion models for high-fidelity long video generation, 2023
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cc329618-d09e-4155-b6d9-3801ce893388 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Classifier-free diffusion guidance, 2022
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a5eca5c-96dc-49aa-8ed2-5b40b5dbfbb3 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Denoising diffu- sion probabilistic models, 2020
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5817e66a-784b-44d6-a5b1-96bb176da9db · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma, Ben Poole, Mohammad Norouzi, David J
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3bdea22d-7a28-4964-ab87-2f792ddb7244 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b7084fb8-1a02-47d7-8ede-a0f0cea68e7f · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality metrics: Psnr vs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4262265-11ba-48d2-82e6-473c74470020 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Dive: Dit-based video generation with enhanced control, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cb2d66d6-ce0b-47b0-97a8-8a025b5a5259 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 423a01fb-3249-47f8-bdad-d8586f9bae21 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Kingma and Jimmy Ba
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 455d1cb4-25aa-493a-b752-b233614df139 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Auto-encoding varia- tional bayes, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f0fb8fa-0c4a-43da-903d-3a7751493cea · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Mpeg: A video compression standard for multimedia applications
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9de07438-a2cb-4fb5-b457-f70e028a892c · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Disentangled motion modeling for video frame interpolation, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 052ce252-8499-4549-a593-102c8f8c9028 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 924c9bcf-4a3a-45d8-a213-923bbb3b3393 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bd3a3ad6-4eda-4771-9c63-771799d62655 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora plan: Open-source large video generation model, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 87ad3afc-bb50-4f2d-b28f-78a0058a8530 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Finite scalar quantization: Vq-vae made simple, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e510d70c-295b-438a-b4cd-be509f36a39e · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Video generation models as world simulators
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 173455d7-a39f-4480-b391-802264b326a1 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Scalable diffusion models with transformers, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad9038f8-7450-447d-81e5-5dc39d72f040 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6643c645-fc39-491d-8589-c90136903690 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics High-resolution image synthesis with latent diffusion models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37570def-13f5-445f-892b-59827a87c91f · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics U-net: Convolutional networks for biomedical image segmentation,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ee7803-ea9c-44c0-abdc-a0cd568b579f · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Adversarial diffusion distillation, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe8ac751-be3d-49fb-970b-7726cbe6aedc · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e5951a-f6f9-429e-b587-3d0f9f4bab8a · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Make-a-video: Text-to-video generation without text-video data, 2022
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5580f9b1-ae07-4cc5-a575-7fd1d4940c33 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Denois- ing diffusion implicit models, 2022
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6b1eed6-8857-47d4-a42f-bab53e265fea · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 85fa5572-3bdd-41d3-ba03-3cfb56585fa4 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics VidTok: A Versatile and Open-Source Video Tokenizer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ea52da-1ee8-453c-b874-64c470e31fc4 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics To- wards accurate generative models of video: A new metric & challenges, 2019
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9876e47-d21f-4e97-a5f6-1c5cacd4a381 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Neural discrete representation learning,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d987cf8c-d8e6-4114-9d4d-f3c0598db836 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca92a608-5a65-4c86-9430-df41fdc0d127 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Phenaki: Variable length video generation from open domain textual description, 2022
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 807e9664-9048-4004-b4f8-8fcc69693e8c · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Mcl-jcv: a jnd-based h
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7de20dd2-15e5-4a32-8ee6-42a7acf73a27 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Bevt: Bert pretraining of video transformers,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9f6493ea-3575-49f0-a9dd-a49a5b66912f · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Emu3: Next-token prediction is all you need, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6290c8bc-e4ca-4e51-b91b-6673e02eda3e · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Instructavatar: Text- guided emotion and motion control for avatar generation,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 795e550f-68be-4473-b774-9977d0adb0a3 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74505f1e-79ca-43f0-bd80-31a878d1698e · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Image quality assessment: from error visibility to structural similarity
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c843c61-c080-42dc-86c6-f7357543a426 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03644ed3-7253-4b81-aa0c-fc0cb8e1b16e · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics ivideogpt: Interactive videogpts are scalable world models, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c4b000f-6b2b-4dc8-8f70-c2d6a6fe5f46 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8fe3a8d6-92aa-4ce5-963c-5cbf606de626 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Cogvideox: Text-to-video diffusion models with an expert transformer, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dda1ea4e-8b0b-4d2f-a6fe-193a372632fc · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Vector-quantized image modeling with improved vqgan, 2022
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3dfeba-01ee-499e-b89e-8fd5b3249106 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4acacae4-5516-4dec-8315-c1db71da2988 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1bcb5cb-a46b-4c0b-9509-f74b5073fe01 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Make your actor talk: Generalizable and high-fidelity lip sync with motion and appearance disentanglement, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1347ab3e-fe86-4e52-8d7b-cd36406f7d22 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Efficient video diffusion models via content-frame motion-latent decomposition, 2024
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dfd38271-7304-4290-8088-3d54a0da7dda · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Efros, Eli Shecht- man, and Oliver Wang
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66641f7-ee8c-470f-b8f1-6105676630a9 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Video in-context learning: Autore- gressive transformers are zero-shot video imitators
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c4fa64eb-468e-42b6-ad53-0923e4ad0126 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Cv-vae: A compatible video vae for latent generative video models,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b91d5f28-1d4c-458a-96e9-693274ea7182 · outbound
VidTwin: Video VAE with Decoupled Structure and Dynamics Open-sora: Democratizing efficient video production for all, 2024
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation abf58d5f-9249-463d-8ac4-8bed29cccf68 · inbound
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion VidTwin: Video VAE with Decoupled Structure and Dynamics
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef828168-849f-4259-8d45-f7f2490571e5 · inbound
Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · inbound
V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.