Pith. sign in

Paper Citation Record · LEDGER

V-RAE: Rethinking Video Latent Spaces for Generation

As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2608.13556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13556 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:07.560939Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1ff04b0-f62e-4edf-8228-4f445fd41fea · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

V-RAE: Rethinking Video Latent Spaces for Generation Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.419865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.419865Z digest=sha256:801fb0a6bec819c83181de22b9c545076219011c2068d5d01eeb1c94d85d1118

Observation 4447f9fa-a3c3-431d-8097-f730f0dfeb96 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.432474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.432474Z digest=sha256:ce7c469c00b61128f50f6f25224fe1ae5bf19b49efa2a8a9a7b846570574044d

Observation bdcfce0e-4edc-4c09-a229-1410ffb3817a · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

V-RAE: Rethinking Video Latent Spaces for Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.437307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.437307Z digest=sha256:5a4e221f7a8eb7ae1bd55aa17880a7fec69672f40302cc550bea9f9e84d0e709

Observation 822e9cf7-499b-4321-8c06-45437b1d7d2c · outbound

This paper cites The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents.

V-RAE: Rethinking Video Latent Spaces for Generation The latent perturbation is used only as a reconstruction-training augmentation; evaluation, latent-statistics estimation, and latent video generation all use clean latents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.235615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T04:16:07.556599Z digest=sha256:d3301ead6d487ad41b1bb2a9ae60a7a9923b6b3006dac2b86a0590bf87da0d56

Observation 7be6db4f-9549-4388-8706-111c5887fa99 · outbound

This paper cites The Kinetics Human Action Video Dataset.

V-RAE: Rethinking Video Latent Spaces for Generation The Kinetics Human Action Video Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.449932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.449932Z digest=sha256:13722e8b38a804da2dc68d629ae43a856187d698c10b334645e3eddba9bc9393

Observation c640daf3-f0a5-4252-9f5c-4bb9c82b55b5 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

V-RAE: Rethinking Video Latent Spaces for Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.453726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.453726Z digest=sha256:31e7b2dac1b9e989cbacd011d80e786778775e98bf71b8abb5f89c7b9ddc6642

Observation 4b88751b-242a-4767-b78d-6090785898a2 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.458705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.458705Z digest=sha256:bd6978ae3a777ed3470d542e63823f530315669f38b6a198eb21bf6531e76a94

Observation 32c377c0-0683-4177-a23f-c8d74756ed76 · outbound

This paper cites Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,.

V-RAE: Rethinking Video Latent Spaces for Generation Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.463799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.463799Z digest=sha256:ff9ceb7da6f75636a0ec5fcac31e8f0f093a1ea31e92a4222a48a7aaf7c1582c

Observation 1715a996-a8eb-4e7c-93d8-7855af826bc1 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.468746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.468746Z digest=sha256:572b19708dd1482cd75cef080396f378108dfc290adf323b1488fe5b8bdb81a7

Observation 24aedaf7-f34f-43c1-988b-43d5371d167e · outbound

This paper cites DINOv3.

V-RAE: Rethinking Video Latent Spaces for Generation DINOv3

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.478488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.478488Z digest=sha256:2626499b50dacab4ee2c76fc0c1992ce7551f140a11edc6de5b125f841e39b3d

Observation bce3e1be-f1a3-454b-b373-e8c803aa75d7 · outbound

This paper cites Improved Baselines with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improved Baselines with Representation Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.483649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.483649Z digest=sha256:64bc8e52b944d8a4ddf22c7818e26418fd97b1b3a45cb0d1b7c72aef5f38a8f5

Observation 4c927a75-48a4-446e-a8db-1c66ec61b5ed · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

V-RAE: Rethinking Video Latent Spaces for Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.492798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.492798Z digest=sha256:515c0038718e63485a2740a920d1fa4e35ab5ad391ecbe185505473c6b060d3f

Observation 1a1d4876-fec9-4f39-b268-392b5243f904 · outbound

This paper cites Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,.

V-RAE: Rethinking Video Latent Spaces for Generation Scaling text-to-image diffusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.501394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.501394Z digest=sha256:2eb50554f23a55dc4a7834534682946665163a67676b230ccf6314572de77a69

Observation dea9f47f-51ca-4dea-94e4-9db511d75c7a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

V-RAE: Rethinking Video Latent Spaces for Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.504838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.504838Z digest=sha256:134f509483f2aec9d8db27d19dd9e905abc0eb698488b6f22e902acf14e8fe2d

Observation 6ed848e9-08e3-4044-99f2-b9f7936e713e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

V-RAE: Rethinking Video Latent Spaces for Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.509532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.509532Z digest=sha256:8af1946db3377bf25ced84977ad47c02bf14b0972b1c0c91959ae01eff4470bb

Observation 0d46f0a2-d88e-4157-87d3-ebf0d5b62093 · outbound

This paper cites Larp: Tokenizing videos with a learned autoregressive generative prior.

V-RAE: Rethinking Video Latent Spaces for Generation Larp: Tokenizing videos with a learned autoregressive generative prior

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.260044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T04:16:07.518767Z digest=sha256:9061b368e51ef092aeebe60a174bfa316f7d9507295b1514a7305f7d15087b71

Observation c0fc0f33-5340-4164-8dc2-b8294f5720c0 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

V-RAE: Rethinking Video Latent Spaces for Generation VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.523981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.523981Z digest=sha256:0179bcddde63770f3b0956fbb8d40a966d8952c671070c9cc66620b538632907

Observation 5c801366-1054-466b-b96e-bbc3f9721aa7 · outbound

This paper cites Making Reconstruction FID Predictive of Diffusion Generation FID.

V-RAE: Rethinking Video Latent Spaces for Generation Making Reconstruction FID Predictive of Diffusion Generation FID

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.532657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.532657Z digest=sha256:466aeac01642d8f99ef3431ec28b58a41e3a02d22f05863f48ffbec6ab83290f

Observation e19e2ac3-fd39-4026-88f7-8aa24b5b4913 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

V-RAE: Rethinking Video Latent Spaces for Generation Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.247634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T04:16:07.536617Z digest=sha256:fe02f062d868260355ad6c40160adf560fb4828c4bbaea901c4007d78d2ed07e

Observation 07ba8789-d4c7-4792-84dd-94679d813625 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

V-RAE: Rethinking Video Latent Spaces for Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.540481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.540481Z digest=sha256:2e27c463a5802789520edd6522dddeb6c5a87b6d8cf0fbe2ac752cc6e0ecace6

Observation 3ce2d571-5fb6-4834-853f-f71141afcf86 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

V-RAE: Rethinking Video Latent Spaces for Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.544422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.544422Z digest=sha256:0252ab705b5a76be3c8df601dfdbcc6274a358cd0b6307d4af9d273629f8dec6

Observation d574185f-1288-415a-aa4d-085b3074756b · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Diffusion Transformers with Representation Autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.548558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.548558Z digest=sha256:88270d33680e7b4f6a3025a00908f6bc373825daf8ff06fb5edabacced0fab9a

Observation 668e9aa1-a0a4-4e1e-bc87-0ea517977c57 · outbound

This paper cites Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,.

V-RAE: Rethinking Video Latent Spaces for Generation Efficient universal perception encoder.arXiv preprint arXiv:2603.22387,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.552563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.552563Z digest=sha256:70eecbc3f8aab8a591fd1bceab9b5a6260ad61a75270fb1a14e5528393294248

Observation f503f54a-b211-4ffd-9eeb-50438244f92e · outbound

This paper cites Only the input channel count and corresponding time shift change for EUPE-B.

V-RAE: Rethinking Video Latent Spaces for Generation Only the input channel count and corresponding time shift change for EUPE-B

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.221888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T04:16:07.560939Z digest=sha256:1d9f93432b43fda4746af86d9690e5feac14516b759946dc2ab21ec8bf9b5573

Observation 4a60b787-5796-4cda-9d61-5ab9045089ec · outbound

This paper cites Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a.

V-RAE: Rethinking Video Latent Spaces for Generation Representation entanglement for generation: Training diffusion transformers is much easier than you think.Advances in Neural Information Processing Systems, 38:7714–7743, 2025a

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.528689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.528689Z digest=sha256:28aa036af2579d9f1d240f3f6253150c35f8fe95d608d983928b3f2ce3e3a17e

Observation caceaa4c-1ffa-4ba9-be4c-a72b7725b96e · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

V-RAE: Rethinking Video Latent Spaces for Generation RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.497419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.497419Z digest=sha256:d23dd3d8d625f18956d431af1f2a95eae50f9dd8fd0d82fa48cd9455b8f4a661

Observation a0fe687e-725d-44d6-b5c2-0b3ce4674630 · outbound

This paper cites Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,.

V-RAE: Rethinking Video Latent Spaces for Generation Dera: Decoupled representation alignment for video tokenization.arXiv preprint arXiv:2512.04483,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.445929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.445929Z digest=sha256:296fed094208bc18f8a1f882d827ff854a0844deff69b04cf69e4176e453d159

Observation ddcbd4b4-db9c-40cf-b329-2bc79b53a10b · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

V-RAE: Rethinking Video Latent Spaces for Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.514157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.514157Z digest=sha256:8eb3baf5730a60336a4c18b2b3c870aa7653c0eb778b99ae6b0b3a2c58b25093

Observation 6d972246-bfff-44ac-bf89-e1863b9d3ed9 · outbound

This paper cites A Short Note about Kinetics-600.

V-RAE: Rethinking Video Latent Spaces for Generation A Short Note about Kinetics-600

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.441723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.441723Z digest=sha256:2a4b6f3d90929a3cf024d9b4771aa8af16dc20003a9e8ea647ad3fffa7570525

Observation 4a781aca-ad0e-49de-a977-ae39dad674cd · outbound

This paper cites V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning.

V-RAE: Rethinking Video Latent Spaces for Generation V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.473562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.473562Z digest=sha256:e364cf211b6441310e4033046fd195185d72df926baf6a8d9075f64ab3af8bf2

Observation ff700b73-2f1e-4f9f-8c12-40d2fd0a3c5c · outbound

This paper cites CoVLA: Comprehensive vision-language-action dataset for autonomous driving.

V-RAE: Rethinking Video Latent Spaces for Generation CoVLA: Comprehensive vision-language-action dataset for autonomous driving

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:16:08.273629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T04:16:07.427895Z digest=sha256:123102ac9476bba3a3bb4584b4d274ff8f028e665efc801272ff06e7e45cfe28

Observation 83a764d7-a9ff-4c8b-a812-46ac2f97e315 · outbound

This paper cites Improving the Diffusability of Autoencoders.

V-RAE: Rethinking Video Latent Spaces for Generation Improving the Diffusability of Autoencoders

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:07.488440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:07.488440Z digest=sha256:befac03328af260fe996e19fcaf67ffc415bdd1dfe97c5c10f4618ddaddf5589

Pith citing papers

No inbound Pith citation observations are available.