Pith. sign in

Paper Citation Record · LEDGER

Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2309.15818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.15818 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:17:24.888752Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T15:07:39.832649Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da46a16a-343c-46e5-934c-826c4a7a0248 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.628714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:3777922543c6e770dc62984c0d7884682dcf9fdf103558cfe4f6e6ff2bad0334

Observation a2eefc06-abbe-4521-bd74-678f58431d37 · inbound

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation cites this paper.

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:34:53.242156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:34:53.141898Z digest=sha256:1e76cc5115a46ac83a439bf1f170ce6655b6f290460c3b297b7a299cf20ddd9f

Observation c8a26450-2e36-4c1b-a2a1-dc8ee2314689 · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:26:22.362066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:8be373daceb7730addafe8a1b66037de8d4fe916a3a6a702ed9fa1804b3283a8

Observation 23c178f8-e3c1-4e2d-a33d-713b0800ab21 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.503525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:d552f1fce24c4ef6b9cdf20e9dd78bd06460c2faf71fc21f4858a09d29f3b440

Observation 31424b6a-0733-4bc8-b814-63af7358e723 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.836921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:a39a8b69e7072fb6bb3d2657c7a62619755ae0bc915a668eb422b8e90b27d3ad

Observation 501f2b6f-5abd-4022-a7bc-d9cacc3cced1 · inbound

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models cites this paper.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.888752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.888752Z digest=sha256:23035c1126ba65aa83b8395aeb24656313feed6bb20069d46081798d62d94a4f

Observation 0dc15b45-3499-4c62-8900-cadfb86ff712 · inbound

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models cites this paper.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.427689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.427689Z digest=sha256:e482260089014b8d3cbf3ce60a22cecc175a4a414435b2ca98cf7e202220e7fc

Observation cb6db24d-d8ec-4485-bcdd-6fcca71cb2a3 · inbound

Enhance-A-Video: Better Generated Video for Free cites this paper.

Enhance-A-Video: Better Generated Video for Free Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:32:46.501520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:32:46.501520Z digest=sha256:6f98177f1661946fedabd09c4b66ae322f621d2edc5b662ca5ce93031ad2e5e1

Observation 79daf691-163a-4b8d-a35a-e57b70bf965e · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.022934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9b75cac18e4f7887f802ebc945e510980706ccabfa57326e76a72ebdc6630640

Observation 10f83b87-7cdb-434c-b36f-e1beda082db1 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.937297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:36b91b1bb5c85217a6601b00eb3dfac0376f02dd516ae78a7edd03fa7e074baa

Observation 11b1ca19-bb6c-4ff8-b3bf-9e08a21dc877 · inbound

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation cites this paper.

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:03.827609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:03:03.827609Z digest=sha256:03d9c8f1d84f8c3dfb56a58a7840cdf5c910c1a222f59eee1409cc0aeddea2e8

Observation edc4c914-eb9e-4cdc-a452-ca4cdb0daf9f · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.153376Z digest=sha256:ea8ee5b83721858e092869c9c45e7bb5ea76c28e682404c83547a4396b3e74f4

Observation 8307ad76-5cfa-48f0-92c3-b7aed07c9196 · inbound

Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model cites this paper.

Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:14:41.760354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:14:41.760354Z digest=sha256:7135b3b6477a0745862185d4db2eed8990c4988ac49de1e48042a91762d404b4

Observation 93bbbced-40a3-4aa5-852d-047e6952a905 · inbound

Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity cites this paper.

Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:52.252058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:31:24.784166Z digest=sha256:38bd29a694b495694c0b98e388b33a62da8e93040b7a0b3ca9804e212ede1f52